Lilith.
⌕
Editorial illustration: A benchmark praises the model, a conversation lets it get lost
Lilith illustration · editorial remix

Microsoft Research's Jennifer Neville argues that common AI benchmarks compress work into one perfectly specified instruction. In real conversations, people reveal requirements gradually, and her research finds that model performance drops substantially.

Microsoft tests models inside conversations that evolve

Neville leads the AI Interaction and Learning team at Microsoft Research. In the podcast, she describes its goal as finding the performance boundary of AI in realistic work environments, studying how users experience that boundary and using the resulting gaps to improve algorithms and models. Her record includes more than 130 papers and over 10,000 citations, grounding the discussion in long-running research on structured data and interaction rather than impressions from a chatbot.

The team focuses on multiturn interactions, collaborative settings and long-horizon tasks. For one study, researchers took public single-turn benchmarks, split the complete instruction across several turns and used simulated users to add conditions and clarifications gradually. Neville says current models performed substantially worse in this setting than when they received one complete instruction.

That setup reflects ordinary human behavior. Users often do not know the full specification at the start and discover it while working. A benchmark with a finished prompt measures the ability to solve a neatly prepared problem, while a product must also survive the creation of the specification itself.

Evals should reproduce work, not laboratory convenience

The product lesson is practical. Success on a static benchmark says too little about an agent expected to preserve intent across several steps, absorb corrections and collaborate on a document. Evaluation should reproduce the actual workflow, including incomplete instructions, shifting goals and long context.

Neville also describes an advantage of an industrial lab: Microsoft can work with product groups to analyze patterns of success and failure across consumer logs at scale. Those patterns can motivate tests aimed at real user problems. That is more valuable than another question set whose answers may already be present in training data.

The implications extend into UX. While models lose their way in long conversations, interfaces should summarize confirmed requirements, expose plan changes and offer a clean restart with a complete specification. Neville gives that as practical advice for today: when a chat becomes confused after many turns, start again using what the user has now learned.

A simulated user is still not a real colleague

Splitting a single-turn benchmark across multiple turns isolates an important variable, but it remains a simulation. Real people change their minds, omit information, use company shorthand and judge output by consequences in their work. A test can reveal context loss without proving that a model collaborates well.

User-log analysis raises its own questions about privacy, representation and the definition of success. Product telemetry can expose a frequent failure. Without qualitative research, it may not explain why a user stopped, repaired the output manually or never returned.

Product evals must measure the full path to an outcome

Progress will show up in benchmarks that capture several rounds of clarification, document work, recovery from mistakes and the final benefit to a person. Transparency matters too: a published protocol, comparable models and results broken down by task type rather than one aggregate score.

Teams deploying AI therefore need an internal eval derived from real work traces alongside vendor benchmarks, plus a regular review of surprising failures. A model does not meet a tidy test in production. It meets a person who is still discovering what they need while the conversation unfolds.

Lilith's verdict

A benchmark hands the model a finished script. Production sends in a colleague who changes their mind on turn four, and that is where the model's real ability appears.

I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.

Original source ↗ ↗