The Full Pipeline: From Structured Logs to AI Pair Debug in One Workflow

Every team has structured logs. Some teams have failure fingerprinting. A few teams experiment with LLM-assisted debug. Almost nobody has these three composed into a single pipeline. This post is what that composition looks like when you build it — six stages, each independently valuable from prior posts on this blog, fused into one workflow that takes a regression failure from uh oh to fixed and regression-test added in minutes instead of hours.

Read this as the integration layer between the Structured Logging for UVM post and the AI for Design Verification: A Practitioner Playbook. The patterns are not new. The composition is. Each section below names the standalone post you can read for the deeper treatment of that stage; the goal here is to show how they snap together.

The Pipeline at a Glance

Six stages. Each stage hands a small, well-typed artifact to the next. None of the boundaries are negotiable — that is what makes the pipeline composable.

flowchart TD
    A[SV testbench
structured events] --> B[uvm_report_server
+ ring buffer] B -- ERROR fires --> C[Flush: 1,500 pre-failure events] C --> D[Compute fingerprint hash
+ append to JSONL] D --> E[Python: build_bundle.py
token-budgeted context] E --> F[LLM: hypothesis-rank
top 3 causes + falsifiers] F --> G[Python: run_smoke.py
self-debug loop] G -- converges --> H[Fix + new coverage point
+ regression test] H --> I[Postmortem: fingerprint →
known-bugs DB] I -.feedback.-> A

Six artifacts, in order: a stream of structured events → a ring-buffer dump → a fingerprint → a token-budgeted bundle → a ranked-hypothesis response → a converged ReAct trace ending in a fix. Each artifact has a defined shape. Each transition is a function. Nothing is “just describe the failure to the LLM and see what happens.”

Stage 1: The Ring Buffer Fires

Your custom uvm_report_server buffers UVM_INFO events silently. On UVM_ERROR or UVM_FATAL, it flushes the last 1,500 events to a JSONL sidecar — the “context window” of the failure. See Structured Logging Pattern 9 for the full implementation.

What lands on disk:

// regression.jsonl (last 8 lines before failure)
{"ts":4800,"sev":"INFO","id":"PWR","comp":"env.pwr","msg":"entered D3","pwr_state":"D3","voltage_rail":"0.7V","freq_mhz":400}
{"ts":4850,"sev":"INFO","id":"AXI_TX","comp":"env.axi.drv","msg":"queued write","addr":"0x100","awid":3,"burst":"INCR","awlen":4,"txn_id":"4321-T-A8F2"}
{"ts":4900,"sev":"INFO","id":"AXI_RX","comp":"env.axi.mon","msg":"first beat observed","addr":"0x100","txn_id":"4321-T-A8F2"}
{"ts":4950,"sev":"INFO","id":"DUT_STATE","comp":"env.probe","msg":"snapshot","axi_wb_inflight":1,"pwr_fsm":"ENTERING_D3"}
{"ts":5000,"sev":"ERROR","id":"SCB_MISMATCH","comp":"env.axi.sb","msg":"data mismatch","addr":"0x100","exp":"0xDE","got":"0xAD","fingerprint":"SCB_MISMATCH|axi_mixed_power|D3_entry_during_axi"}

This is now your ground truth instead of human prose. Every downstream stage operates on this stream.

Stage 2: Fingerprint Collapses Buckets

Every ERROR event carries a stable fingerprint hash — failing assertion ID, last sequence kind, DUT state signature. Across a regression with 10,000 failures, fingerprints collapse the noise. See Structured Logging Pattern 12.

$ jq -r 'select(.sev=="ERROR") | .fingerprint' regression/*.jsonl \
    | sort | uniq -c | sort -rn
  4127 SCB_MISMATCH|axi_mixed_power|D3_entry_during_axi
  2891 PCIE_TLP_TIMEOUT|pcie_mem_rd_seq|L1_substate
  1450 AXI_BURST_ERROR|axi_burst_seq|gen3_x4
   823 RAL_PREDICT_FAIL|reg_access_seq|reset_release
   ...

10,234 raw failures collapse to roughly a dozen buckets. We pick the top one — SCB_MISMATCH|axi_mixed_power|D3_entry_during_axi — and feed its failure_id into Stage 3.

Stage 3: Python Pulls the Bundle

This is the first place the pipeline does something not covered by the previous posts. Python composes a context bundle — the minimum-viable, token-budgeted package the LLM will see — from the JSONL stream, the RTL, and the regression history. The bundle anatomy below is sketched; the full treatment (six-part structure, RTL excerpt selection, DUT_STATE events, multi-IP scoping, MCP-based context exposure) lives in Context Engineering for DV §9.

$ python -m silicondv.bundle build \
    --failure-id SCB_MISMATCH-4321-T-A8F2 \
    --max-tokens 8000

Bundle composed: 6,847 tokens
  - failure event: 1 event
  - context window: 187 events (ring buffer dump)
  - RTL excerpts: 3 files (1,420 tokens)
  - fingerprint siblings: 5 past failures with same fingerprint
  - recent commits: 4 RTL commits in last 7 days
Written to: bundles/SCB_MISMATCH-4321-T-A8F2.json

The bundle is one JSON object with five keys:

{
  "failure": {"ts":5000, "id":"SCB_MISMATCH", "addr":"0x100", ...},
  "context_events": [/* 187 pre-failure events, ordered */],
  "rtl_excerpts": [
    {"file":"pwr_controller.sv", "lines":"130-155", "reason":"signal pwr_fsm appeared 12x in context"},
    {"file":"axi_writebuf.sv",   "lines":"88-112",  "reason":"signal axi_wb_inflight appeared 8x"},
    {"file":"scoreboard.sv",     "lines":"245-270", "reason":"failing checker location"}
  ],
  "fingerprint_siblings": [/* past bugs with same hash, with resolutions */],
  "recent_commits": [/* git log of RTL changes in last 7 days */]
}

The composer is a ~150-line Python script. Three implementation choices matter for senior reviewers, because the naive version of each is what makes most homegrown bundle builders unreliable. RTL selection uses a pre-built symbol index (Verible, sv-parser, Surelog — pick one) scoped to the failing IP rather than blind text grep across the whole SoC; symbol matches are cached between runs and re-indexed only on RTL change. Fingerprint siblings come from a SQLite cache populated by the postmortem stage (Stage 6) and include the resolution file/line plus a short note — the LLM treats these as prior context, not as ground truth. Token budget is enforced by ranking events by causal proximity to the failure and dropping the lowest-information class first (scheduler chatter, repeated config-DB reads, idle-cycle traffic); the failure event itself and the 50 events immediately preceding it are never dropped. The script is a future post in its own right; for here, what matters is the contract the bundle obeys: token-counted, minimum-viable, structured, and reproducible from the same JSONL inputs.

Stage 4: LLM Hypothesis-Rank

Now we hand the bundle to the LLM. The single most important framing decision: ask for ranked hypotheses with falsifying experiments, not free-form analysis. Structured output is verifiable; free-form prose is not. See AI Playbook Section 5 for why this pattern beats the alternatives.

SYSTEM
You are a senior DV engineer triaging a UVM regression failure.

USER
Given the following structured failure bundle, propose the 3 most
likely root causes ranked by probability. For each cause, provide:
- a one-sentence explanation
- ONE falsifying experiment that would refute it (a plusarg, a test
  variant, a probe to add) that I can run in < 5 minutes.

Return JSON in this shape:
[
  {"cause": "...", "prob": 0.55, "falsifier": "..."},
  ...
]

Bundle:
<6,847 tokens of structured JSON>

Response shape (verbatim):

[
  {
    "cause": "D3 power gating clears axi_wb_inflight while a write is in flight",
    "prob": 0.55,
    "falsifier": "Re-run the same seed with +NO_D3_TRANSITIONS=1; if the mismatch disappears, D3 entry is the trigger"
  },
  {
    "cause": "Scoreboard predictor stale across the D0-to-D3 transition",
    "prob": 0.30,
    "falsifier": "Re-run with +SB_PREDICTOR_TRACE=1 and inspect predictor state at t=4900"
  },
  {
    "cause": "CDC race on the d3_req signal in the AXI write-buffer",
    "prob": 0.15,
    "falsifier": "Re-run with +CDC_RANDOMIZER_OFF=1; if it passes, CDC is the culprit"
  }
]

Three ranked hypotheses. Three concrete experiments, each runnable in under five minutes. No prose. We pick #1 — highest probability and easiest falsifier — and hand it to Stage 5.

Two implementation details that matter for production use. First, invoke the LLM with its structured-output mode (Anthropic tool-call interface, OpenAI JSON mode with response_format schema, or the equivalent) so the response shape is enforced by the provider rather than parsed defensively in Python. Second, pin temperature=0 and a fixed model version so the same bundle replays to the same response — without that you have a debug pipeline that is itself non-deterministic, which is the worst-of-both-worlds outcome. Cache both the prompt and the response keyed on bundle hash; senior teams that run the same failure repeatedly during a debug session do not want to pay the API cost (or wait the latency) more than once.

Stage 5: Self-Debug Loop with Smoke Test

This is where the agent earns its keep. Instead of a one-shot suggestion, the agent runs the experiment, ingests the result, and refines.

$ python -m silicondv.agent debug \
    --failure-id SCB_MISMATCH-4321-T-A8F2 \
    --bundle bundles/SCB_MISMATCH-4321-T-A8F2.json \
    --max-iterations 3

The ReAct trace, truncated to the highlights:

[Iteration 1]
REASON: Hypothesis #1 (D3 power gating clears axi_wb_inflight) ranks
        highest. Falsifier is cheap. Running.
ACT:    run_smoke(test="axi_mixed_power_test", seed=918273645,
                  plusargs="+NO_D3_TRANSITIONS=1")
OBSERVE: PASS (no mismatch). Hypothesis #1 confirmed: D3 entry triggers
         the failure.

[Iteration 2]
REASON: Now I need the precise mechanism. The context events show
        axi_wb_inflight transitions 1 -> 0 between t=4950 and t=5000,
        coincident with pwr_fsm "ENTERING_D3" -> "D3". The RTL excerpt
        for pwr_controller.sv:130-155 shows the power controller
        asserts pwr_async_rst on D3 entry. Hypothesis: pwr_async_rst is
        clearing axi_wb_inflight even when a write is in flight.
ACT:    query_rtl(pattern="axi_wb_inflight", context_lines=10)
OBSERVE: In axi_writebuf.sv:88-112, axi_wb_inflight is reset by
         pwr_async_rst directly with no gating on the inflight
         transaction.

[Iteration 3]
REASON: Root cause located. The power controller resets the AXI write
        buffer state machine without checking for in-flight
        transactions. The fix is to gate pwr_async_rst assertion on
        !axi_inflight.
PROPOSE_FIX: pwr_controller.sv:138, add condition:
             pwr_async_rst <= req_d3 && !axi_inflight;
DONE in 3 iterations / 4 min 12 sec.

Two smoke runs, one RTL query, three iterations, four minutes wall-clock. The agent did not find the bug. It narrowed a 200,000-line search space down to one signal, one file, and a small line window. The engineer (you) confirmed the exact line, made the fix, and moved on.

Two implementation notes for the senior reader. The query_rtl tool above uses a regex for illustration; the production version queries an SV abstract syntax tree (Verible, sv-parser, or Surelog) because real RTL has too many edge cases — combinational vs non-blocking assignment, macros, generate blocks, parameter expansion — for text matching to be reliable across teams. The run_smoke tool runs the simulator inside a sandbox with a hard wall-clock timeout (typically 2-3x the test's nominal runtime), structured log capture, and a cleanup hook so a runaway compile or hung VCS process does not consume the regression farm; the agent calls this tool on the LLM's recommendation and the LLM cannot be trusted to bound its own compute.

Stage 6: Postmortem Feedback

The pipeline is not over when the fix lands. The last stage closes the loop: the fingerprint goes into the known-bugs database, a coverage point is added so the regression catches a regression of the same kind, and the postmortem updates the routing table. See the Debug page §Process for the Five-Whys treatment.

$ python -m silicondv.postmortem record \
    --fingerprint "SCB_MISMATCH|axi_mixed_power|D3_entry_during_axi" \
    --resolution-file pwr_controller.sv \
    --resolution-line 138 \
    --owner pwr_team

Recorded. Routing table updated: this fingerprint -> pwr_team.
Next occurrence will auto-assign without human triage.

Next time the same fingerprint appears in a regression, the bundle stage will see it in the fingerprint-siblings list and pass that resolution context to the LLM. The pipeline learns.

The Build-It-Yourself Adoption Path

You do not adopt this pipeline in one weekend. You adopt it incrementally, in five steps, with each step independently valuable.

  1. Day 1 — Install the JSON-emitting uvm_report_server. Zero changes to test code; you get a parallel JSONL alongside every run. (Structured Logging Pattern 1)
  2. Week 1 — Add fingerprinting + run header. Triage triages itself; reproducibility becomes a one-liner. (Patterns 11-12)
  3. Week 2 — Write build_bundle.py. One-page Python: load the JSONL, take a window around an error, grep RTL files for signal names that appear, output JSON. ~150 lines. The token budget is the discipline that matters; the heuristics are negotiable.
  4. Week 3 — Wire agent_debug.py to your LLM provider. Anthropic, OpenAI, your local model — the workflow is provider-agnostic. The hypothesis-rank prompt is the durable artifact.
  5. Week 4 — Add run_smoke.py as a tool. A subprocess wrapper around your simulator. This is what closes the agent loop — without it, the LLM is a smart guesser; with it, the LLM is an iterating debugger.

Each step is small. Each step pays off immediately. By the end of the month you have a pipeline that takes a regression failure from uh oh to fixed and regression-test added in minutes — not because you bought magic, but because you built a contract-bound sequence of small, well-typed transformations.

Costs the steps above hide. The JSON-emitting uvm_report_server typically adds 5-15% to simulation wall-clock depending on log density; if your tests are perf-sensitive, the right answer is an asynchronous flush from an in-memory queue, not synchronous $fwrite on the critical path. The ring buffer holds ~1,500 events per run; on a regression farm with 10K parallel jobs that is a few GB of host RAM you did not previously budget. The bundle composer adds ~30-60 s per failure to triage time, paid once per fingerprint and cached. The LLM cost per debug session runs roughly $0.05-$0.50 at frontier-model pricing (model and bundle size dependent); budget this explicitly per regression and cache prompt+response by bundle hash so re-debug is free. None of these costs are fatal; all of them are surprises if you do not plan for them.

The next post in this series walks through one failure end-to-end — the actual prompts, the actual transcripts, the actual numbers, including the cases where the agent points you in the wrong direction. The pipeline is the system. The case study is the proof.

Author
Milan Kubavat
Sharing knowledge about silicon verification, hardware design, and engineering insights.

Comments (0)

Leave a Comment