Mechanistic interpretability, done from outside the field and published with its null results
I study what a language model does on the inside. There is an original paper submitted to BlackboxNLP 2026, the first consolidated Spanish-language course on the subject, and a new project published before it has results. I come from eighteen years in energy, and that is where this points.
A model's thought about a fact splits into parts — the thing, and the question asked of it — and those parts suffice: average them, recombine them, write them back inside, and the model speaks the answer. Three models, two domains where it works and two where it provably doesn't. Interactive demo included.
The first consolidated course on the subject in Spanish: six pages from "what is a token" to attribution graphs. Alongside it, a map of the six pillars of AI safety, ten open research lines, and resources curated link by link.
The lab notebook: 989 lines in chronological order with the experiments, the measurement bugs I hit while building, and the replication that died under matched controls and turned into the paper. Every number recomputes from the artifacts.
The newest one, still without results: residual-gauge — importing a basis for the internal state from a symmetry the task already has, with thresholds pre-registered before anything ran.
Where this is going
AI safety will reach energy. It hasn't yet
Frontier-model use outside sanctioned controls is already measured inside critical-infrastructure organizations. Once agents move from advising to acting on a pipeline or a grid, someone will have to show when that is safe. I spent eighteen years inside that industry, and it strikes me as where this ends up.
What I cannot claim: nothing I research today applies to an operating decision. My experiments separate "Italy" from "capital-of" in the internal state of a small model, in English, over toy factual relations. Between that and auditing an agent with authority over a pipeline there is a gap I would rather state than paper over.
The six pillars of AI safety, where interpretability sits among them, and the four rungs still missing to close that gap: the full map (in Spanish).
Who's behind this
Matías Podeley. Eighteen years in energy in Argentina — reservoir engineering, then business development. Now independent interpretability research. Buenos Aires.
If you work in this
Funders and collaborators
If you fund research in this area, run a fellowship or incubator, or work in AI safety and want a collaborator who knows energy operations from the inside — email me.