Hoppa till innehåll
VibekollenBETAVibekollen
VideoAI Engineer

Your Agent Evolved. Your Evals Didn't. — Ameya Bhatawdekar, Braintrust

Teams built orchestration graphs because the models of 2024 could not be trusted to orchestrate, and then the models learned to orchestrate and the graphs became the thing holding them back.

Ameya Bhatawdekar traces that loop across five generations of architecture, each one forced by a step change in model capability, and argues that evals have to move with it. A single prompt needed only answer quality. A retrieval chain added a parser that grabs the wrong field and a retriever that returns the wrong context. Graphs added branch logic, contracts between nodes, and classifier nodes that misfire quietly, which is a great deal of new surface to check. What changed most recently is not another layer but the unit of measurement. Once a loop is reliable enough to run free, the same input produces visibly different trajectories on every run, so a single eval result stops meaning very much. He separates the two questions it hides. Pass at k asks whether the system succeeds at least once across k attempts, which measures capability. The stricter variant asks how many of those k attempts succeed, which measures reliability.

Öppna på YouTube →

Sammanfattningen är skriven av Vibekollen utifrån källans egen publicering. Innehållet tillhör AI Engineer.

Mer från AI Engineer