Writing
Notes on LLMs, interpretability and building AI products by Sai Anirudh Siddi.
2026
Grading an LLM is harder than training one
Anyone can run an RL loop now. Knowing if the model actually got better is the hard part, and it comes down to reward hacking, LLM judges, rubrics and benchmarks that are sometimes wrong themselves.
The crossroads: local LLM or cloud API
Cloud models are faster and smarter, local models keep your data yours. Thoughts on the tradeoff from someone stuck at that crossroads, and why privacy isn't the argument.
Persona-vector probes for medical LLMs
Reading harmfulness, uncertainty, and other behaviors straight off a medical LLM's activations, and the custom probe builder that lets a clinician define new ones on demand.
From 0.50 to 0.88 AUROC with activation normalization
Why normalizing residual-stream activations before a linear probe is the highest-leverage change for behavioral detection in LLMs.
Sparse autoencoders, live
An SAE turns a model's opaque residual stream into a cloud of readable features. Here's how GlassBox uses one as a live interpretability view, and why we don't trust it just on its own.
2025
No posts match that yet. Show all posts.