Ambient Intelligence

Cloud models are too slow for real-time human interaction. We examine tiny interaction models, sub-10ms on-device sensory loops, and asynchronous cloud planner architectures.

Attention: Undivided & Uncompressed

Nearly ten years after the "Attention Is All You Need" paper, attention still powers the most capable LLMs in the world. We take a look at how attention has evolved in the last decade, from its kernel geometry to the silicon it runs on.