← Back to Blog
Process

Evidence Over Claims: Building Durable Artifacts

One of the easiest ways to build false confidence in agentic systems is to trust agent summaries over actual output. An agent might claim "I ran the tests and they passed," but did it really? ADLC solves this by demanding that every stage of work produces durable, observable artifacts that humans can inspect.

The Problem: Trusting Summaries

When an agent completes a task, it typically returns a human-readable summary: "I wrote the feature, added tests, and verified that it works." But this summary is just a claim. It's not evidence.

What actually happened?

  • Did the tests actually pass, or did the agent skip slow tests?
  • Are the added tests meaningful, or just smoke tests?
  • What code was written? Did it introduce new dependencies or security issues?
  • What assumptions did the planning agent make that the implementation agent didn't catch?

Without looking at the actual artifacts, you're flying blind.

Seven Types of Artifacts

ADLC captures evidence at every stage:

  1. Plans: Not just the ticket, but the research behind it. What viability checks did the planner run? What are the explicit assumptions and dependencies? Are there flagged risks?
  2. Context: What code, docs, and workspace history did the agent reference? This helps future agents understand what sources were trusted.
  3. Run logs: Full stdout/stderr from the agent's execution, including all tool calls, timestamps, and any errors or warnings that occurred.
  4. Code diff: The exact changes proposed, with line-by-line diff format so reviewers can see what was added, removed, or modified.
  5. Test output: Full test runner output, including pass/fail status, coverage percentages, timing data, and any skipped or flaky tests.
  6. Review notes: Comments from automated checkers (lint, security scan, type check) and notes from human reviewers.
  7. Delivery evidence: The PR/MR link, merge status, deployment timestamp, and post-deployment health checks.

Why Artifacts Matter

These artifacts serve multiple purposes:

  • Review: Humans can make informed decisions based on facts, not abstracts. They can question assumptions and spot issues before they ship.
  • Debugging: When something goes wrong, the artifacts are a complete record of what the agents did. No guessing or re-running tasks.
  • Learning: Future agents can search and analyze these artifacts. "What similar problems have we solved?" "Why did this assumption fail last time?"
  • Compliance: For regulated industries, these artifacts are the audit trail. Every decision is documented with its supporting evidence.
  • Optimization: Tracking latency, cost, and failure patterns across hundreds of tasks reveals where to improve the system.

Structuring Artifacts for Query

ADLC doesn't just dump logs into a bucket. Artifacts are structured:

  • Metadata: Task ID, timestamps, agent type, model used, token counts, cost, and outcome status.
  • Full-text indexed: Plan descriptions, code comments, test names, error messages are searchable by future agents and humans.
  • Linked: Each stage references the previous stages, so tracing a decision back to its origin is straightforward.
  • Retention policy: Artifacts follow configurable retention (keep forever for compliance, or archive after 1 year for cost).

This structure means that when a future planner asks "Has anyone tried adding WebAssembly support?", the system can search all historical plans and show what was learned.

What NOT to Rely On

Some outputs are useful but not authoritative:

  • Agent summaries: Used for human readability, but not for determining success or failure.
  • Telemetry gaps: If a test didn't run, that's not evidence that everything is fine—it's evidence that testing was incomplete.
  • Agent confidence scores: An agent claiming 95% confidence means nothing. Show the evidence that backs it up.

Implementing Evidence Capture

In practice, this means:

  • Every tool call is logged with input, output, and any errors.
  • Test runs save the full output, not just a pass/fail bit.
  • Diffs are stored in their original format (unified diff) plus parsed into a queryable format.
  • Reviews capture not just decisions but the reasoning for each decision.
  • All artifacts are immutable once the stage completes (prevent retroactive claim changes).

Read the full brief: ADLC Brief
Back to blog: All posts