• TurnBench: Grading Turn-Taking Like a Linguist, Not a Timer

    A short, figure-led breakdown of TurnBench: a multi-domain, conversation-analysis-grounded benchmark for end-of-turn and interruption detection in spoken dialogue.

  • Your eval is a measuring instrument

    Notes on building evals worth trusting: traces over questions, rubric before instrument, quality before cost, and failure names specific enough to test.

  • Where Does a System Prompt Go? Taking PersonaPlex Apart

    NVIDIA's PersonaPlex adds role conditioning and zero-shot voice cloning to a full-duplex speech model without changing the architecture. A two-person walkthrough of the Hybrid System Prompt, the synthetic data pipeline, the six-hour fine-tune, and Service-Duplex-Bench.

  • The User Is Not a Turn: Omni-Flow and MiniCPM-o 4.5

    MiniCPM-o 4.5 treats user speech as part of the observed environment rather than a conversational turn. A two-person walkthrough of Omni-Flow's chunked serialization, why one-second chunks beat 100 ms, TAIL's time-aligned interleaving, and the four-stage training pipeline.

  • Taking Moshi Apart: A Conversation About Every Piece

    A two-person walkthrough of every component of Moshi — Helium, the Mimi codec, the RQ-Transformer, multi-stream modelling, Inner Monologue, the four-stage training ladder, quantization, safety and evaluation — covering what each piece did, why, what it changed, and what it leaves open.

  • Evolving Voices: How Speech Became a Language Modeling Problem

    A walk through the architectures that turned continuous audio into something a language model can predict — wav2vec 2.0, HuBERT, WavLM, SoundStream, AudioLM, VALL-E.

  • The Gatekeeper of Voice AI: A Deep Dive into Voice Activity Detection

    From Likelihood Ratio Tests in WebRTC to self-attention in Transformers, the algorithmic evolution of Voice Activity Detection.

  • Who Plays the User? The State of Full-Duplex Evaluation

    A dozen benchmarks now measure full-duplex speech models and they disagree about almost everything. The reason is one shared unsolved problem: who plays the user.

  • How Audio Becomes Tokens

    Codecs, RVQ and the Mimi tokenizer, explained from first principles: how a 24 kHz waveform and a stream of words end up as integers on the same 12.5 Hz clock.

  • How Full-Duplex Audio Models Actually Work

    Moshi, dGSLM, PersonaPlex, Lychee-FD and NVIDIA VoiceChat 11B, explained from first principles: why voice AI still feels like a walkie-talkie, and what it takes to fix it.