COSCUP x UbuCon Asia 2026 標誌
  • 關於我們
  • 議程
  • 交通
  • 會場地圖
  • 社群
  • 贊助夥伴
    • 參與指南
    • 第一次參與
    • 活動參與
    • 講者參與
    • 前夜派對
    • 海外參與者
    • 開源社群、攤位及議程軌
    • 贊助夥伴
    • 邀請函申請指南
  • 工作人員
  • 周邊活動 / BoF
  • 部落格
  • 社群守則
English

AI 也要期中考? The Story Behind Building a Benchmark for Agent Workflow Development

時間
2026-08-09 10:05 ~ 10:35
講者
Kent Huang
位置
AU
共筆
Open LLM End User: 開源模型應用 入門中文

議程簡介

要考期中考?AI 沒過也會有這天。開發 Agent Workflow 時要怎麼分辨。表現變好是 Model 的功勞,還是好的 Workflow 帶來的功勞?這時就需要一套好的驗證方法:期中考!但是道高一尺,魔高一丈。你想得到的,AI 也想得到。上網作弊、偷看答案,只有你想不到,沒有 AI 做不到。本節的故事背景,是我們在開發開源 Agent Workflow 時,設計內部 benchmark 所碰到的各種大小事。

How do you evaluate the performance of an agent workflow? When things get better, is it the model doing the work, or the workflow itself? A good benchmark is like a midterm exam. It tells you who actually studied. But AI is smarter than you expect. It will cheat by searching the web, peek at other answers, and find every loophole you forgot to close. AI always finds a way to surprise you. In this session, we'll share the story behind building a benchmark for an open-source agent workflow. Every twist, every workaround, and every time the AI quietly outsmarted our test design.

講者

Kent Huang

Kent Huang

Software Engineer @ Recce 與 AI 搏鬥中,嘗試不要被淹沒在 AI 洪流的工程師一枚

鑽石級

Canonical

黃金級

連續贊助2 年中華民國資訊經理人協會連續贊助2 年國泰金融控股公司連續贊助5 年玉山銀行累計贊助12 年MySQL累計贊助6 年華捷智能股份有限公司Nitra

白銀級

累計贊助16 年慧邦科技股份有限公司  Gamesofa Inc.

青銅級

華娛網路娛樂股份有限公司財團法人國家實驗研究院國家高速網路與計算中心ONLYOFFICEQNAP Systems, Inc. 威聯通科技連續贊助16 年祐生研究基金會源鋼技術顧問有限公司 SUN SQUARE Co., Ltd

好朋友級

連續贊助3 年晶心科技股份有限公司沛星互動科技股份有限公司

特別感謝

連續贊助2 年臺北市政府資訊局美商賽發馥股份有限公司臺灣分公司Rozeta AI

共同主辦單位

累計合作9 年臺灣科技大學 電子工程系

協辦單位

累計合作12 年開放文化基金會

COSCUP x UbuCon Asia 2026

Conference for Open Source Coders, Users, and Promoters | 亞洲最大開源年會,由社群自發舉辦。

聯絡我們

  • 會眾服務
  • 贊助合作
  • 議程投稿
  • 行銷方案

相關資源

  • COSCUP 部落格
  • 訂閱電子報
  • 活動照片

網站導覽

  • 首頁
  • 關於我們
  • 交通資訊
  • 贊助夥伴
20062007200820092010201120122013201420152016201720182019202020212022202320242025