Mechanistic interpretability and applications: The final fronteir of hacking
- Time
- 2026-08-08 12:30 ~ 13:00
- Speaker
- Martin Chang
- Room
- TR212
- Co-write
Abstract
LLM changed the course of tech and humanity for better or for worse. But safety has and still is a problem. No one can guarantee if LLMs are evil or misaligned. Instead of trying to make LLMs safe during training by different means. Mechanistic interpretability provides a different way - try to detect and change LLM behavior by opening them up and see what is going on.
Mechanistic interpretability is an important tech and not an easy one. Yet they are rarely discussed or developed publicly. Let's change that. If not so future us down the road can control much more powerful AI then right now.
Speaker
Martin Chang
Systems software engineer working on HPC, GPGPU, and AI.