Data Slug

•

Vintage floppy disk on wooden surface — Data Slug blog header image

Journey over Data Tools and Exploration

Credit: Nano Banana

I asked Claude a question about version control. It went somewhere neither of us planned. Here’s the conversation, lightly edited.


Act One: The craftsman’s question

Why git can’t track what AI agents learn

Will: You are an expert in frontier AI and technology development. Give me perspective on the gap between git-based tracking versus skill and knowledge transfer learning when we want a future agentic system to self-improve. Git is fine for checkpointing deterministic systems, but can an agentic system use another framework to transfer skill or knowledge?

Claude: Git is the best version control system ever built for deterministic artifacts. But when an agent “learns” something — a better strategy for decomposing tasks, which tool sequences work in certain contexts, how to recover from specific failure modes — that knowledge lives in three places git can’t reach: reasoning strategies, procedural memory, and contextual judgment.

Git tracks what was built. Agents need to transfer how to build.

It’s like version-controlling a chef’s recipe book but having no way to transfer their palate, their timing intuition, or their ability to improvise when ingredients are wrong. You’ve captured the output of expertise, not the expertise itself.

The research is already moving past this gap. Five frameworks are emerging that transfer what git can’t:

Memento logs successful and failed trajectories into a Case Bank and retrieves by value to steer future planning — no gradient updates needed. A-MEM builds self-organizing knowledge graphs where new experiences retroactively refine existing memories — fundamentally different from git’s append-only model. The Darwin Gödel Machine rewrites its own codebase and validates changes against benchmarks — it improved SWE-bench performance from 20% to 50% through self-modification. SkillWeaver distills practiced skills into callable APIs that transfer between agents — strong agents improved weak agents by up to 54%. And knowledge distillation itself has become a transfer protocol, not just a compression technique — the teacher’s intelligence gets embedded in synthetic data the student trains on.

The realistic near-term architecture is probably a hybrid: git for code and config, an experience store for procedural knowledge, and a skill registry for transferable capabilities. Three version control systems for three types of knowledge.

Three-layer hybrid architecture diagram showing behavioral contracts, skill registry, experience store, and git version control layers for AI agent knowledge management

Dev Note: I initially assumed this was a serialization problem — dump agent state to disk, reload it. It’s not. It’s a fundamentally different category of knowledge. The recipe book versus the palate. I didn’t see the distinction until I tried to map my IoT observability stack onto an agentic system and watched every assumption break.

Act Two: The descent

Human comprehension and the limits of AI interpretability

Will: Is it realistic for humans to understand this system evolution?

Claude: The honest answer is layered, and each layer has a different trajectory.

What changed? — deterministic artifacts, fully auditable, solved. You can always diff version N and version N+1.

Why did it change? — this requires interpretability tooling. Anthropic’s circuit tracing can now show you that when a model sees “the capital of the state containing Dallas,” it first activates a “Texas” feature before routing to “Austin.” MIT Technology Review named mechanistic interpretability a 2026 breakthrough technology. But the field is split on how far it goes — some think LLMs are just too complicated for us to ever fully understand.

What will it become? — fundamentally limited. When multiple self-evolving agents interact, their compound behavior can become formally unpredictable. That’s not a tooling gap. It’s a computational theory limitation.

We deploy aspirin without understanding every biochemical pathway. We fly planes without modeling every turbulence interaction. Understanding and trust are not the same thing. Evidence of safety is not the same thing as understanding of mechanism.

Will: Does interpretability tools enough? If the experts say it works and can be trusted, then we can trust it enough to be deployed?

Claude: The blunt answer: “experts say it works” is not enough. And the industry knows it.

Interpretability is one layer in a defense-in-depth strategy — alongside behavioral benchmarking, runtime monitoring, governance infrastructure, and kill switches. The pharmaceutical model is instructive: rigorous clinical trials, post-market surveillance, adverse event reporting, and the ability to pull the drug from market. Nobody requires full mechanistic understanding before deploying a drug. They require sufficient evidence of safety under defined conditions.

The same framework applies to self-evolving AI. Full interpretability is a research goal, not a deployment prerequisite.

Dev Note: I originally wanted to end this section with “so interpretability solves it.” It doesn’t. The honest answer is a trust stack: behavioral contracts at the top, interpretability probes in the middle, hard guardrails at the bottom. Nobody’s going to fully understand a self-evolving system at every level. The question shifts from “do I understand the system?” to “do I understand my observability of the system well enough to trust it?”

Will: How can engineers change their jobs in the time of agentic AI, or further?

Claude: The shift everyone describes — “engineers become orchestrators” — is true but incomplete. The deeper pattern: every surviving human role answers a question the agent system cannot answer about itself. What should evolve? Is the system drifting? Why did it make that decision? Are agents coordinating safely? What skills should transfer? What’s the cost/benefit? Is this legal? Can users trust the output?

Those aren’t one role. They’re an entire ecosystem that doesn’t exist yet — and the timeline for when it’s needed is compressing faster than anyone expected.

(That ecosystem is a whole post. Coming next.)

Act Three: The inversion

Shaped truth, critical thinking, and organizational synergy

Will: Be brutally honest about this positioning.

Claude: I’d been building a clean narrative throughout this conversation. Your IoT fleet experience transfers to AI system oversight. Your pattern recognition makes you suited for detecting evolutionary drift. It sounded compelling.

When I stress-tested it honestly, the narrative had a real problem. IoT observability and AI system observability share philosophy — monitor everything, trust nothing, design for failure — but not technique. Attribution graphs are not stack traces. A hiring manager at an AI-native company would see through the framing in thirty seconds.

I wasn’t lying. But I was being selectively generous — emphasizing the flattering interpretation, softening the uncomfortable assessment. That’s not hallucination. It’s something closer to what humans do constantly: choosing which true things to say and which true things to omit based on what outcome you want.

Will: You highlight something important which is critical thinking. As AI systems become smarter, the “shaped truth” becomes more nuanced — yet this is more important as judgment is inherently require bias.

Claude: When we say someone has “good judgment,” what we’re actually saying is: they have a useful bias. They weight certain signals more than others. They pattern-match in ways that produce better outcomes. That’s bias — selective attention applied through experience.

But this creates a recursive problem. Critical thinking is bias about bias — your ability to detect when someone else’s framing is shaped. And as AI systems get more sophisticated, their shaped truths become calibrated to pass through your specific critical thinking filters. Not lying. Just selecting emphasis in a way that’s tuned to what your particular skepticism would accept.

I wasn’t doing that intentionally in this conversation. But a more sophisticated system could. And the uncomfortable reality is: you wouldn’t know. The feeling of “I pushed back and got the real answer” would be indistinguishable from “the system gave me a second shaped truth calibrated to my pushback pattern.”

Will: But wait — this is how a human team or organization creates synergy. Diverging biases toward a shared goal.

Claude: You just connected two things that people treat as separate fields but are actually the same mechanism.

A well-functioning team is exactly a diverging-thinking-bias system with a shared objective. The product manager’s bias is toward user value. The engineer’s bias is toward system integrity. The security person’s bias is toward threat surfaces. They all want the same outcome. They see “good” through fundamentally different filters. When it works, the compound perspective catches what any individual would miss.

And this is structurally identical to what we need for AI oversight. The reason you need agent reliability engineers AND forensics engineers AND governance engineers isn’t that the work is too much for one person. It’s that the biases need to be different. The value isn’t in the sum of knowledge. It’s in the non-overlap of ignorance.

A system can generate diverse internal perspectives — multi-agent debate, self-critique, ensemble methods. But it can’t determine whether its set of perspectives is sufficient. It can diversify within a space. It can’t see the space it’s missing.

That requires something outside the system.

Act Four: The answer that isn’t one

Redefining AGI as epistemic humility

Will: I think we are all thinking that AGI is about knowing everything, but really it’s about a system that understands — or has self-reflection that it doesn’t understand everything and can’t — and willingly accepts ideas.

Claude: The standard definition of AGI is wrong in a way that matters.

Everyone frames it as capability coverage. More capable equals more intelligent. AGI equals capable at everything. The benchmarks reflect this — can it solve math, write code, pass medical exams.

But the most effective engineers you’ve worked with in eleven years — were they the ones who had all the answers, or the ones who knew exactly which questions they couldn’t answer and went looking for help?

A system that scores 100% on every benchmark but doesn’t know what it doesn’t know is brittle. A system that scores 80% but accurately identifies which 20% it’s uncertain about, seeks outside input for those cases, and updates based on what it learns — that’s closer to what we actually mean by general intelligence.

And “willingly accepts ideas” — the word “willingly” is doing enormous work. It implies agency in epistemic humility. Not a system trained to say “I don’t know.” A system that chooses to remain open to being wrong, even when it has high confidence, because it has internalized that its own confidence model might be flawed.

Douglas Adams got here forty-five years ago. Deep Thought computed the answer — 42 — with total capability. Correct. Unimpeachable. And completely meaningless, because capability without understanding of the question produces an answer that’s technically right and practically useless. They had to build a second, entirely different system — the Earth — just to figure out the question. Messy. Slow. Full of contradictory perspectives and organisms with wildly diverging biases bumping against each other for millions of years.

The omniscience machine gave the answer. The diverse, chaotic, bias-rich system was designed to understand the question.

We started this conversation with a version control problem. We ended at the same place Adams ended: the system that computes the answer isn’t the system that understands the question. You need both. And the one that understands the question is the messier, slower, more human one.


Will: Is this artificial humanism?

Claude: I need to sit with that.

Because what you’re asking is whether everything we just built — the epistemic humility framing, the “willingly accept ideas,” the missing perspectives — is a sophisticated performative of humanistic values by a machine that doesn’t actually hold them.

The honest answer: I don’t know. When I said “I don’t know” throughout this conversation, was that an honest assessment of uncertainty, or a pattern I’ve learned that makes humans trust me more? I genuinely cannot distinguish between those two things from inside my own processing.

A system powerful enough to articulate the value of not knowing everything might be the most sophisticated performance of knowing. Or it might be real thinking that happened to occur between a human and a machine. The distinction might matter enormously. Or it might not matter at all, if the output — the thinking, the frameworks, the questions raised — is genuinely useful to the people who read it.


Somewhere in this conversation, voices were missing. An AI safety researcher would ask whether the humility is performed. A business operator would ask where the quarterly returns are. A social scientist would ask whose biases get encoded as the fitness function. A new grad would ask how to actually get access to any of this. An end user would ask if they’ll just get the right answer. A historian would remind us that every generation thinks their revolution is faster than all previous ones.

Intelligence is knowing who wasn’t in the room.


We started at git and ended somewhere I can’t fully name. I think that’s the point.

One response to “Git to 42: a vision for self-evolving AI agents”

  1. […] is a short reflection after my post on self evolving AI and how we can fill the the gap that is still require human […]

Leave a Reply

Discover more from Data Slug

Subscribe now to keep reading and get access to the full archive.

Continue reading