The Story Behind Building a Benchmark for Agent Workflow Development
- Time
- 2026-08-09 10:05 ~ 10:35
- Speaker
- Kent Huang
- Room
- AU
- Co-write
Abstract
How do you evaluate the performance of an agent workflow? When things get better, is it the model doing the work, or the workflow itself? A good benchmark is like a midterm exam. It tells you who actually studied. But AI is smarter than you expect. It will cheat by searching the web, peek at other answers, and find every loophole you forgot to close. AI always finds a way to surprise you. In this session, we'll share the story behind building a benchmark for an open-source agent workflow. Every twist, every workaround, and every time the AI quietly outsmarted our test design.
Speaker
Kent Huang
Software Engineer @ Recce 與 AI 搏鬥中,嘗試不要被淹沒在 AI 洪流的工程師一枚