davidfacer.com / aimaturitymodels.com / AI-Native Maturity Models / AI-Native SDLC / D12

Instrumentation & observability

When an AI agent produces a piece of generated code, a test, or an infrastructure change, and something later goes wrong with it, can anyone actually reconstruct why the agent produced what it did — or is the only evidence the artifact itself, with no trace of the reasoning behind it? This dimension measures whether an organization can answer that question, and whether the answer scales as generation volume grows.

Where most organizations start (Nascent)

Monitoring covers infrastructure and application metrics, but organizational knowledge — decisions, plans, the reasoning behind generated artifacts — lives in human-readable formats that aren't structured for agent consumption at all. The first real step is simply retaining traces of build-time generation activity: which agent or model was involved, the relevant inputs and outputs, tool activity, and lineage — enough to actually investigate a failure instead of relying on memory.

Where the real gains happen (Modeled → Integral)

The meaningful shift is standardizing that retention and using it as real evaluation evidence — identifying recurring failure categories and feeding findings back into how specifications, generation, architecture, and testing actually work, rather than treating each trace as one-off debugging material. The bigger gain beyond that is recognizing that AI can participate in more than just generation: identifying every place it participates in the delivery system's own runtime decisions — self-provisioning, live release gating — and instrumenting those separately from build-time generation, since they're genuinely different surfaces.

What the top of the curve actually looks like (Telemetric)

At full maturity, build-time traces, delivery-system runtime traces, and production telemetry are all equally and directly accessible to agents and humans alike, feeding back into the upstream shared intelligence layer (D1–D3) as part of one closed loop. The organization keeps re-evaluating where non-determinism actually lives in its own delivery system as the other dimensions mature, rather than treating any past answer as permanent.

Why this dimension matters

D12 is what makes D13 possible — it supplies the evidence of the return; feedback loop velocity (D13) measures how fast that evidence actually gets used. Without this dimension, "learning from what happened" stays a slogan instead of a mechanism.


Drafted from the SDLC model's real locked content — including the transition and verification notes now folded into ai_native_sdlc_maturity_model.md itself (2026-07-27). D12 carries no open review flag.

Drafted from the SDLC model’s real locked content.

Feed This to Your AI© 2026 David Facer