# Context graphs are more ambitious than they sound

13 January 2026

A concept is gaining traction: context graphs as the missing enterprise layer. Not a system of record for what happened, but for why it happened.

The idea resonates because the gap is real. AI is moving from advising to acting. Once systems start acting, "why" stops being philosophical and becomes operational. Money moves. Access changes. Promises get made. The organization is bound, and "why" is what makes those commitments legible and defensible.

The pitch sounds right. But it hides an assumption that will hurt you later: that "why" will emerge if you store enough traces.

Why doesn't emerge. It has to be bound.

## The cut: discovery vs commitment

Before getting to what context graphs get wrong, it helps to draw a line most discussions blur.

**Discovery mode** is where AI shines: retrieval, comparison, synthesis, hypothesis, simulation. You explore possibilities and build understanding. A context graph can be extraordinary here: surfacing precedents, finding patterns, answering "what have we done before?"

**Commitment mode** is where organizations become bound: spend, access, price, promise, change. The moment output turns into a real-world move that the business has to live with.

The danger is treating discovery artifacts (traces, patterns, "what usually happens") as permission to commit.

Inference is for insight. Commitment needs reference.

Draw that line and the hype gets easier to evaluate.

## "Context graphs will capture the why"

When people say a context graph will "capture the why," they're usually mixing two different questions:

**Justification:** Why was this action allowed, and what would have forced a stop?

**Causality:** Why did this action work? Did it cause the outcome? Under what conditions would it fail next time?

Traces can suggest both. They can't certify either.

A context graph built from post-commit exhaust reconstructs what happened. It shows sequences, correlations, patterns. It offers plausible narratives. What it can't do is certify that the narrative was the binding interpretation that made the commitment legitimate at the time.

Consider a concrete case. The exhaust shows a leader repeatedly rejecting contracts from a vendor over eighteen months. A context graph surfaces this pattern. An agent infers "reject this vendor" and starts doing so automatically.

But the real rule was "reject until they achieve ISO 27001 certification." The vendor gets certified in month nineteen. The inferred rule is now wrong, and the system keeps rejecting while sounding perfectly consistent. Your competitor closes the deal. Your team spends weeks trying to explain why the system "doesn't trust" a qualified vendor.

The condition was never recorded. The trace showed correlation. It didn't show the rule.

Any time the rule is conditional and the condition isn't in the graph, precedent turns into superstition.

## "We can infer the rules from patterns"

The escape route is familiar: we don't need to capture the rules explicitly. With enough examples, the system will learn what's allowed.

This is where the idea stops being harmless. It confuses pattern recognition with operational authority.

Patterns tell you what usually happens. They don't tell you what's allowed to happen. The model can't distinguish between "this was correctly approved under the rules" and "this happened and nobody caught the error." Both look identical in the trace.

Worse, meaning drifts. "Approved" meant one thing when the delegation policy was written. It means something different after three reorgs. "Strategic account" had a definition in 2022. It has a different definition now. The graph stores artifacts that use the same words with shifting referents. Query it, and you get a weighted average of contradictory meanings.

Inference without reference is guesswork at scale. And guesswork becomes policy the moment it can move money.

If you can't point to something checkable (which rule-set was in force, what evidence was admissible, what authority applied) you don't have decision memory. You have story memory. And story memory is why post-hoc compliance feels like archaeology.

## "This is premature, agents aren't deployed yet"

Organizations are still struggling with data unification. Agents aren't operating at scale. Instrumenting decision traces now feels like building for a future that hasn't arrived.

The objection makes sense if the only path is "agents everywhere generating traces."

But the conclusion changes if you start at the commit boundary.

You don't need agents everywhere. You need clarity at the moments the organization becomes bound. Those moments exist today, with humans in the loop, inside existing systems: the approval that moves money, the access grant that opens data, the exception that changes terms.

Start there. Require that when a question closes, the system records what made the answer valid, and what would change it.

This doesn't require a perfect data lake. It doesn't require agents. It requires binding the regime when the question closes: which rules were in force, what evidence counted, and whose authority applied, fixed before anything acts on it.

Immediate value, even before automation:

- Fewer unauthorized commitments slip through
- Faster approvals with less rework
- Decisions that survive employee turnover
- Reuse that stays safe when policies drift

The decision anchors the trace. Without it, you're building a context graph that can tell you what happened but not whether it was allowed.

## What has to be true

If justification is going to be real, it has to be captured when the question closes, before anything acts on it, not reconstructed after.

When a question closes, the system should be able to say why the answer was valid, and what would have forced a stop, escalation, or expiry.

Those answers define the regime in force when the question closed. Pin them, and the graph becomes trustworthy. Skip them, and you're storing exhaust and calling it memory.

One subtlety matters here: identity tells you who clicked. It doesn't tell you whether they had standing to bind the organization at that moment, under the regime in force. People retain identity long after authority changes. If you don't pin what was decided, and on whose authority, when the question closes, traces preserve the appearance of legitimacy after the basis has shifted.

A permission says who could act. It doesn't say what was decided. Context graphs that blur this distinction will confidently retrieve precedents that no longer apply.

## The reuse problem

Reuse is the prize everyone wants. A queryable record of what worked before, so you don't start from scratch every time.

But reuse without structure is silent replay.

"We did it this way last time" is dangerous unless you know what had to be true for that precedent to apply. The reuse check has two parts.

First, does the earlier answer still hold? Same question? Same definitions in force? Same authority regime? Did any of the answer-changing facts change?

Second, is it enough for this use? The earlier decision was established on a particular basis. A new case can ask more of it, even when nothing has changed: more evidence, a higher approval, a stricter bar.

If both hold, reuse safely. If not, you need to know exactly what shifted, or what is still missing.

Similarity is not applicability. Reuse is a condition check against the prior decision's stated terms, not a similarity search.

Most systems can't do that check because they never captured the conditions. That's why precedent doesn't compound. It replays until something breaks.

## What context graphs are good for

None of this is an argument against context graphs. It's an argument about where they sit.

Context graphs are powerful for discovery and navigation. Surface precedents. Find patterns. Answer "what have we done before?" Run simulations: what-ifs, policy sandboxes, training runs that don't grant real authority.

They become safe for commitment and automation only when anchored to what was decided, and under which regime. Otherwise you're simulating on shifting meanings and calling the output "learning."

If context graphs are the memory layer, decisions are the anchors that make the memory trustworthy.

---

Context graphs are likely inevitable. Organizations want a queryable record of experience. Agents will make retrieval and synthesis cheaper. "Precedent becomes searchable" is too useful to ignore.

But "capturing the why" is more ambitious than it sounds.

The operational "why" isn't a property of logs. It's the allowance structure: what made the action allowed, under which definitions, under which authority, with which evidence, within which bounds. That doesn't emerge from exhaust. It has to be bound when the question closes, before anything acts on it: on the binding path, not beside it.

Trust isn't queryable history. Trust is an action you can trace to a decision you can defend.

---

References:

Jaya Gupta & Ashu Garg ( Foundation Capital ), [AI’s Trillion Dollar Opportunity: Context Graphs.](https://foundationcapital.com/context-graphs-ais-trillion-dollar-opportunity/)

Jaya Gupta, [Where Context Graphs Materialize.](https://www.linkedin.com/pulse/where-context-graphs-materialize-jaya-gupta-lsqoe/)

Related:

Animesh Koratana, PlayerZero, [How to build a context graph.](https://www.linkedin.com/pulse/how-build-context-graph-animesh-koratana-6abve/)

Dharmesh Shah, Context Graphs: [The Elegant Idea Everyone's Talking About](https://simple.ai/p/what-are-context-graphs)
