Towards Trustworthy AI: Essential Open-Source Tools for Agent Bias Detection and Mitigation
- Time
- 2026-08-08 11:30 ~ 12:00
- Speaker
- Jing-Tian Sung
- Room
- TR411
- Co-write
Abstract
As LLM-driven Autonomous Agents become increasingly integrated into automated workflows—ranging from resume screening and customer service to strategic business decisions—their autonomy brings both efficiency and the hidden risk of social bias and discrimination. Because Agents possess the ability to reason and invoke external tools, their biases are often more covert and "actionable" than those of standard LLMs. Evaluating these agents effectively has become a critical challenge in building Trustworthy AI.
In this session, we will dive deep into an open-source bias detection toolkit for Agents. We will deconstruct the technical roots of agentic bias and introduce how to build automated benchmarks using custom knowledge bases and datasets to quantitatively evaluate biased behaviors in specific scenarios.
Key highlights of this session include:
Dimensions of Bias Identification: Defining bias patterns and benchmark adjustments within specific application contexts.
Evaluation Framework Architecture: How to leverage open-source frameworks to build automated testing scripts that simulate diverse user inputs to trigger potential biases.
Real-world Case Studies: Showcasing detection results in practical scenarios, such as recruitment agents or decision-support systems.
Mitigation Strategies: Exploring ways to reduce bias risks during the development phase through Prompt Engineering and Guardrails.
Through these open-source tools, we aim to empower developers to proactively monitor and rectify Agent behavior. This is more than a technical showcase; it is a call to the open-source community to collaborate in building a firewall of fairness and transparency for the AI era.
Speaker
Jing-Tian Sung
試圖在冰冷的模型參數中,找回技術應有人文溫度的工程師一枚。不只是在校正誤差,更想修復的是科技與社會之間的信任鏈。