AI Pair Debug, Worked: An AXI Burst Mismatch in D3 Power State

This is a worked case study. A failing test, walked end-to-end through the pipeline from The Full Pipeline post. The prompts are real verbatim. The LLM responses are illustrative but representative of what a frontier model produces given this bundle shape. The numbers at the end are real measurements from running this workflow on similar bugs in a working DV environment.

Read this as evidence the workflow described in the previous post actually works — and an honest accounting of where it does not. If you want the patterns abstracted, read the previous post first. If you want to see what the pattern feels like in practice, this is that.

Note: the specific bug below is a constructed scenario — an AXI burst mismatch during a D3 power state entry. It was chosen because it is realistic, has clear before/after stages, exercises three different subsystems (AXI, power, scoreboard), and avoids any IP or customer disclosure. The pipeline patterns and the LLM behavior are real.

The Setup

Test: axi_mixed_power_test, seed 918273645. A randomized sequence of AXI write bursts interleaved with periodic power-state transitions between D0 (active) and D3 (deep sleep). Lasts roughly 1.2 million cycles.

Failure: One scoreboard mismatch at simulation time 5,000 ns. Expected data 0xDE, observed 0xAD, at address 0x100. Single failure in an overnight regression of 800 tests — rare enough to be hard to dismiss, frequent enough that letting it slip is not an option.

What the structured log shows (the last 8 events before the error, from the ring-buffer dump):

{"ts":4800,"sev":"INFO","id":"PWR_STATE","comp":"env.pwr","msg":"entered D3","pwr_state":"D3","voltage_rail":"0.7V","freq_mhz":400}
{"ts":4850,"sev":"INFO","id":"AXI_TX","comp":"env.axi.drv","msg":"queued write","addr":"0x100","awid":3,"burst":"INCR","awlen":4,"data_first":"0xDE","txn_id":"4321-T-A8F2"}
{"ts":4900,"sev":"INFO","id":"AXI_RX","comp":"env.axi.mon","msg":"first beat observed on bus","addr":"0x100","awid":3,"txn_id":"4321-T-A8F2"}
{"ts":4920,"sev":"INFO","id":"DUT_STATE","comp":"env.probe","msg":"snapshot","axi_wb_inflight":1,"pwr_fsm":"ENTERING_D3","axi_wb_addr":"0x100"}
{"ts":4960,"sev":"INFO","id":"DUT_STATE","comp":"env.probe","msg":"snapshot","axi_wb_inflight":0,"pwr_fsm":"D3","axi_wb_addr":"0x000"}
{"ts":4980,"sev":"INFO","id":"AXI_RX","comp":"env.axi.mon","msg":"completion observed","addr":"0x100","bresp":"OKAY","txn_id":"4321-T-A8F2"}
{"ts":5000,"sev":"ERROR","id":"SCB_MISMATCH","comp":"env.axi.sb","msg":"data mismatch","addr":"0x100","exp":"0xDE","got":"0xAD","txn_id":"4321-T-A8F2","fingerprint":"SCB_MISMATCH|axi_mixed_power|D3_entry_during_axi"}

Traditional triage estimate: The senior engineer assigned this would open the FSDB, hunt for axi_wb_inflight transitions around 4920 ns, correlate with the power FSM trace, walk the AXI write-buffer module, and probably find this in 3-4 hours. We are going to do it in under 30 minutes.

First Attempt: The Naive Paste

Before we use the pipeline, let us show what happens with the “just paste the log” approach — the move every team starts with and that the AI Playbook warns against.

USER (verbatim)
Here is a failing UVM regression log from an AXI test. What is wrong?

[paste of last 2,000 lines of human log, ~38,000 tokens]

The response (excerpted; the full thing is ~600 tokens of mostly hedging):

ASSISTANT
Looking at the log, several things stand out:

1. There is an AXI write that the scoreboard reports as a mismatch.
   This could be:
   - A timing issue on the AXI bus
   - A race condition in the driver or monitor
   - A scoreboard predictor issue
   - An issue with the burst encoding

2. The error occurs near a power state transition. This could be
   unrelated, but worth checking.

3. I would recommend enabling more verbose AXI tracing and checking
   for protocol violations...

To narrow this down further, could you share:
- The relevant AXI driver code
- The scoreboard implementation
- The exact RTL module being tested
- The simulator version

Generic. Hedging. Asks for four follow-up artifacts the engineer already has in front of them. Mentions the power transition as “could be unrelated” — missing that it is in fact the trigger. This is what an unstructured paste produces from any frontier model: the context-window-degradation effect (arxiv 2509.21361) plus the lack of a focused prompt frame combine to produce a response that is not wrong, but is not useful.

Lesson: the model is not going to compensate for your context discipline. You have to bring it.

Second Attempt: With a Context Bundle

Now we use the pipeline. First, build the bundle (the bundle's structure, ordering rules, and selection heuristics are deep-treated in Context Engineering for DV §9):

$ python -m silicondv.bundle build \
    --failure-id SCB_MISMATCH-4321-T-A8F2 \
    --max-tokens 8000

Bundle composed: 6,847 tokens
  - failure event: 1 event
  - context window: 187 events (ring buffer dump)
  - RTL excerpts: 3 files (1,420 tokens)
      * pwr_controller.sv:130-155 (signal pwr_fsm hit 12x)
      * axi_writebuf.sv:88-112    (signal axi_wb_inflight hit 8x)
      * scoreboard.sv:245-270     (failing checker location)
  - fingerprint siblings: 0 past failures with this fingerprint
  - recent commits: 4 RTL commits in last 7 days
Written to: bundles/SCB_MISMATCH-4321-T-A8F2.json

6,847 tokens — an order of magnitude smaller than the naive paste. Now the prompt:

SYSTEM
You are a senior DV engineer triaging a UVM regression failure.

USER
Analyze this structured failure bundle and identify the most likely
root cause. Be specific about which signal, file, and timing window
you suspect.

[bundle, 6,847 tokens]

Response (excerpted):

ASSISTANT
The failure is highly correlated with a D3 power state entry. From the
context events:

- At t=4800, the system enters D3 (voltage_rail=0.7V, freq=400MHz).
- At t=4850, an AXI write to 0x100 is queued.
- At t=4920, the DUT_STATE probe shows axi_wb_inflight=1, pwr_fsm=
  ENTERING_D3.
- At t=4960, axi_wb_inflight is unexpectedly 0 while pwr_fsm=D3 —
  the write is still in flight on the bus but the internal flag has
  been cleared.
- At t=4980, the AXI completion observes on the bus.
- At t=5000, the scoreboard reports a mismatch.

My primary hypothesis: the power controller is asynchronously clearing
the AXI write-buffer state during D3 entry, even though a write is
mid-flight on the bus. This would explain the corrupted data.

Looking at the RTL excerpt for pwr_controller.sv:130-155, the
pwr_async_rst signal is asserted on D3 entry. I would investigate
whether this signal has an unconditional path to the AXI write
buffer's state machine reset.

Recommendation: examine pwr_controller.sv around the pwr_async_rst
assertion logic, and axi_writebuf.sv around the inflight-flag reset
in its always_ff block.

A genuinely useful response. Names the right subsystems (power controller, AXI write buffer). Identifies the relevant signal (pwr_async_rst). Points at the right files within 4-10 lines of the actual fix location. Still: it is one hypothesis, asserted as the primary, with no falsifier. A senior engineer wants verification before touching RTL. We tighten the frame.

Iteration 1: Hypothesis-Rank

We re-prompt with the structured-output discipline from the AI Playbook:

USER
Thank you. Reframe your analysis as the 3 most likely root causes
ranked by probability. For each cause, provide ONE falsifying
experiment I can run in <5 minutes (a plusarg, test variant, or
probe).

Return JSON in this shape:
[
  {"cause": "...", "prob": 0.XX, "falsifier": "..."},
  ...
]

Response:

[
  {
    "cause": "Power controller's pwr_async_rst clears axi_wb_inflight without checking for in-flight transactions",
    "prob": 0.55,
    "falsifier": "Re-run the same seed with +NO_D3_TRANSITIONS=1; if the mismatch disappears, D3 entry is the trigger"
  },
  {
    "cause": "Scoreboard predictor is stale across the D0-to-D3 transition and reports a stale expected value",
    "prob": 0.30,
    "falsifier": "Re-run with +SB_PREDICTOR_TRACE=1 and inspect predictor state at t=4900; if the expected value differs from 0xDE, the predictor is the issue"
  },
  {
    "cause": "CDC race on d3_req signal between the power clock domain and the AXI clock domain",
    "prob": 0.15,
    "falsifier": "Re-run with +CDC_RANDOMIZER_OFF=1; if it passes, CDC is the culprit"
  }
]

Three ranked hypotheses. Three concrete experiments, each runnable in under five minutes. Hypothesis #1 is highest probability and has the cheapest falsifier (one plusarg flip). We start there.

Iteration 2: Self-Debug Loop

We invoke the agent to run the experiments and converge.

$ python -m silicondv.agent debug \
    --failure-id SCB_MISMATCH-4321-T-A8F2 \
    --bundle bundles/SCB_MISMATCH-4321-T-A8F2.json \
    --max-iterations 3

The ReAct trace (excerpted to the highlights):

[Iteration 1]
REASON  Hypothesis #1 ranks highest. The falsifier is one plusarg.
        Running.
ACT     run_smoke(test="axi_mixed_power_test", seed=918273645,
                  plusargs="+NO_D3_TRANSITIONS=1")
OBSERVE PASS in 47 seconds. No mismatch. D3 entry confirmed as the
        trigger. Hypothesis #1 confirmed; #2 and #3 not yet falsified
        but lower priority.

[Iteration 2]
REASON  D3 entry is the trigger. Now I need the precise mechanism.
        The DUT_STATE events at t=4920 and t=4960 show axi_wb_inflight
        transitions from 1 to 0 between them, coincident with
        pwr_fsm transitioning ENTERING_D3 to D3. From the RTL excerpt
        for pwr_controller.sv:130-155, pwr_async_rst is asserted on
        the D3 transition. I need to see how axi_wb_inflight is reset
        in the write-buffer module.
ACT     query_rtl(pattern="axi_wb_inflight\\s*<=", context_lines=12)
OBSERVE In axi_writebuf.sv around line 92:
            always_ff @(posedge clk or posedge pwr_async_rst) begin
              if (pwr_async_rst) axi_wb_inflight <= 1'b0;
              else if (axi_aw_handshake) axi_wb_inflight <= 1'b1;
              else if (axi_b_handshake)  axi_wb_inflight <= 1'b0;
            end
        The pwr_async_rst is in the asynchronous reset list with no
        gating on in-flight state. Confirmed root cause.

[Iteration 3]
REASON  Root cause located. The power controller's reset clears the
        write-buffer state without regard for in-flight transactions.
        The data the AXI monitor observes (0xAD) is whatever was on
        the bus after the reset corrupted the buffer; the scoreboard's
        expected value (0xDE) is what was queued before the reset.
PROPOSE_FIX  In pwr_controller.sv near line 138, gate the assertion
             of pwr_async_rst on absence of in-flight AXI activity:
             pwr_async_rst <= req_d3 && !axi_inflight_top;
             Alternatively, in axi_writebuf.sv, remove pwr_async_rst
             from the always_ff sensitivity list and replace with a
             synchronous, conditional reset that respects the
             handshake. The first fix is smaller; the second is more
             defensively designed.

DONE in 3 iterations / 4 min 12 sec.
Proposed fix saved to: fixes/SCB_MISMATCH-4321-T-A8F2.diff

Two smoke runs, one RTL query, three iterations, four minutes. The agent did not invent the bug; it narrowed a 200,000-line search space down to one signal, one file, and a 4-line window. It even proposed two alternative fixes with a brief design judgment.

The Fix and Validation

The engineer (you) opens pwr_controller.sv and looks at the proposed line. The agent said line 138; the actual fix turns out to be at line 134 — the agent was off by a few lines but in the right neighborhood. The fix is a one-line change:

// Before
pwr_async_rst <= req_d3;

// After  
pwr_async_rst <= req_d3 && !axi_inflight_top;

Smoke test passes. Full regression re-runs in 30 minutes — passes. A new coverage point is added so the regression catches a regression of the same kind:

// In the power coverage group
cover (d3_entry_req && axi_inflight_top) {
  bins suppressed_during_axi = {1};
  bins fired_during_axi      = {0};  // should never increment
}

And the postmortem is recorded:

$ python -m silicondv.postmortem record \
    --fingerprint "SCB_MISMATCH|axi_mixed_power|D3_entry_during_axi" \
    --resolution-file pwr_controller.sv \
    --resolution-line 134 \
    --owner pwr_team \
    --notes "pwr_async_rst gated on !axi_inflight_top"

Recorded. Routing table updated.

Next time this fingerprint appears, the bundle stage will surface the resolution in the fingerprint-siblings list, and the LLM will incorporate it as prior context.

What Worked, What Did Not

An honest accounting matters more than a victory lap. Here is what landed and what did not.

What worked

  • The bundle stage rescued the LLM. The same model produced near-useless prose from a raw log paste and a sharp, well-aimed hypothesis from a 6,847-token bundle. The context-engineering principle is real and measurable.
  • Hypothesis-rank produced a falsifiable starting point. Asking for ranked hypotheses with falsifiers (rather than analysis) forced structured output that we could act on without thinking too hard.
  • The self-debug loop converged in 3 iterations. Two smoke runs were enough to confirm the hypothesis and locate the root cause to a 4-line RTL window.
  • The RTL excerpt selector got lucky and useful. The heuristic that picks RTL by signal-name hit count put pwr_controller.sv and axi_writebuf.sv in front of the LLM — without that, the model would have asked us to paste the code.
  • Reproducibility was built in. Every LLM call used temperature=0 and a pinned model version, so the same bundle replays to the same response; the full agent transcript was logged with bundle hash, prompt, response, and tool-call sequence so the session is auditable after the fact. For regulated flows (automotive, medical, aerospace) this is non-negotiable; for everyone else it is what makes the session debuggable when the agent goes off the rails.

What did not

  • The naive prompt was a waste. Pasting the raw log produced a generic, hedging response. Without the bundle, the pipeline does not function. Anyone trying to skip steps 3-5 will conclude the LLM “does not work for DV.”
  • 2 of 3 hypotheses were wrong. The rank is a starting point, not a verdict. If hypothesis #1 had been wrong, we would have spent another smoke run on #2. Treat the ranks as priorities for the engineer, not as truth from the model.
  • The proposed RTL line number was off by ~4 lines. Frontier models are not reliable at exact line-number precision. The agent gave us the neighborhood (within a 25-line excerpt); we found the exact line. This is the right division of labor — do not expect more.
  • The model proposed two fixes; both were viable but only one was idiomatic for this team. Architectural judgment on which fix to take is still ours. The agent narrowed the choice; we made it.

Time Comparison

Measured end-to-end times for this category of bug, from one engineer running similar workflows on similar bugs over a regression season.

StepTraditionalThis workflow
Identify failure cluster30 min (read logs by hand)10 sec (jq + uniq -c)
Build mental model of failure45 min (waveform navigation)30 sec (build_bundle)
Hypothesize root cause30 min (think + sketch)2 min (LLM hypothesis-rank)
Verify hypothesis30 min (manual probe + re-run)6 min (2 smoke runs at ~3 min each)
Locate exact RTL line30 min (grep + read)2 min (agent's RTL query + human verification)
Apply fix + validate30 min10 min
Total triage time~3 hours~25 minutes

Not in the comparison: the one-time pipeline setup cost (estimated 1-2 weeks for a working JSON-emitting uvm_report_server, bundle composer, and agent harness). The full regression re-run (~30 minutes, but it happens in either workflow). Code review on the fix (~15 minutes, again the same). The LLM API cost for this debug session (~$0.10-$0.30 at frontier-model pricing — trivial in absolute terms but worth budgeting at scale). The ~5-15% simulation overhead the JSON-emitting report server adds during regression, and the few GB of host RAM consumed by ring buffers across a parallel regression farm; both are real and both have mitigation patterns (async flush, smaller ring depth for non-debug runs) that you implement once and forget.

The 25-minute number is door-to-door triage time after the pipeline is in place. Your mileage will vary: some bugs are easier (the agent solves them in 5 minutes; some are harder (the agent points you in the wrong direction and you fall back to traditional methods). The honest average across a season is roughly 10x faster than the traditional baseline — less for easy bugs, more for those where the LLM context-engineering rewards your investment.

The workflow, not the model, is the durable thing. Swap Claude for GPT-5 for Gemini for your local model — the bundle, the hypothesis-rank, the self-debug loop, the smoke gate are model-agnostic. Pick your provider. Commit to the workflow.

The Full Pipeline: From Structured Logs to AI Pair Debug in One Workflow

Every team has structured logs. Some teams have failure fingerprinting. A few teams experiment with LLM-assisted debug. Almost nobody has these three composed into a single pipeline. This post is what that composition looks like when you build it — six stages, each independently valuable from prior posts on this blog, fused into one workflow that takes a regression failure from uh oh to fixed and regression-test added in minutes instead of hours.

Read this as the integration layer between the Structured Logging for UVM post and the AI for Design Verification: A Practitioner Playbook. The patterns are not new. The composition is. Each section below names the standalone post you can read for the deeper treatment of that stage; the goal here is to show how they snap together.

The Pipeline at a Glance

Six stages. Each stage hands a small, well-typed artifact to the next. None of the boundaries are negotiable — that is what makes the pipeline composable.

flowchart TD
    A[SV testbench
structured events] --> B[uvm_report_server
+ ring buffer] B -- ERROR fires --> C[Flush: 1,500 pre-failure events] C --> D[Compute fingerprint hash
+ append to JSONL] D --> E[Python: build_bundle.py
token-budgeted context] E --> F[LLM: hypothesis-rank
top 3 causes + falsifiers] F --> G[Python: run_smoke.py
self-debug loop] G -- converges --> H[Fix + new coverage point
+ regression test] H --> I[Postmortem: fingerprint →
known-bugs DB] I -.feedback.-> A

Six artifacts, in order: a stream of structured events → a ring-buffer dump → a fingerprint → a token-budgeted bundle → a ranked-hypothesis response → a converged ReAct trace ending in a fix. Each artifact has a defined shape. Each transition is a function. Nothing is “just describe the failure to the LLM and see what happens.”

Stage 1: The Ring Buffer Fires

Your custom uvm_report_server buffers UVM_INFO events silently. On UVM_ERROR or UVM_FATAL, it flushes the last 1,500 events to a JSONL sidecar — the “context window” of the failure. See Structured Logging Pattern 9 for the full implementation.

What lands on disk:

// regression.jsonl (last 8 lines before failure)
{"ts":4800,"sev":"INFO","id":"PWR","comp":"env.pwr","msg":"entered D3","pwr_state":"D3","voltage_rail":"0.7V","freq_mhz":400}
{"ts":4850,"sev":"INFO","id":"AXI_TX","comp":"env.axi.drv","msg":"queued write","addr":"0x100","awid":3,"burst":"INCR","awlen":4,"txn_id":"4321-T-A8F2"}
{"ts":4900,"sev":"INFO","id":"AXI_RX","comp":"env.axi.mon","msg":"first beat observed","addr":"0x100","txn_id":"4321-T-A8F2"}
{"ts":4950,"sev":"INFO","id":"DUT_STATE","comp":"env.probe","msg":"snapshot","axi_wb_inflight":1,"pwr_fsm":"ENTERING_D3"}
{"ts":5000,"sev":"ERROR","id":"SCB_MISMATCH","comp":"env.axi.sb","msg":"data mismatch","addr":"0x100","exp":"0xDE","got":"0xAD","fingerprint":"SCB_MISMATCH|axi_mixed_power|D3_entry_during_axi"}

This is now your ground truth instead of human prose. Every downstream stage operates on this stream.

Stage 2: Fingerprint Collapses Buckets

Every ERROR event carries a stable fingerprint hash — failing assertion ID, last sequence kind, DUT state signature. Across a regression with 10,000 failures, fingerprints collapse the noise. See Structured Logging Pattern 12.

$ jq -r 'select(.sev=="ERROR") | .fingerprint' regression/*.jsonl \
    | sort | uniq -c | sort -rn
  4127 SCB_MISMATCH|axi_mixed_power|D3_entry_during_axi
  2891 PCIE_TLP_TIMEOUT|pcie_mem_rd_seq|L1_substate
  1450 AXI_BURST_ERROR|axi_burst_seq|gen3_x4
   823 RAL_PREDICT_FAIL|reg_access_seq|reset_release
   ...

10,234 raw failures collapse to roughly a dozen buckets. We pick the top one — SCB_MISMATCH|axi_mixed_power|D3_entry_during_axi — and feed its failure_id into Stage 3.

Stage 3: Python Pulls the Bundle

This is the first place the pipeline does something not covered by the previous posts. Python composes a context bundle — the minimum-viable, token-budgeted package the LLM will see — from the JSONL stream, the RTL, and the regression history. The bundle anatomy below is sketched; the full treatment (six-part structure, RTL excerpt selection, DUT_STATE events, multi-IP scoping, MCP-based context exposure) lives in Context Engineering for DV §9.

$ python -m silicondv.bundle build \
    --failure-id SCB_MISMATCH-4321-T-A8F2 \
    --max-tokens 8000

Bundle composed: 6,847 tokens
  - failure event: 1 event
  - context window: 187 events (ring buffer dump)
  - RTL excerpts: 3 files (1,420 tokens)
  - fingerprint siblings: 5 past failures with same fingerprint
  - recent commits: 4 RTL commits in last 7 days
Written to: bundles/SCB_MISMATCH-4321-T-A8F2.json

The bundle is one JSON object with five keys:

{
  "failure": {"ts":5000, "id":"SCB_MISMATCH", "addr":"0x100", ...},
  "context_events": [/* 187 pre-failure events, ordered */],
  "rtl_excerpts": [
    {"file":"pwr_controller.sv", "lines":"130-155", "reason":"signal pwr_fsm appeared 12x in context"},
    {"file":"axi_writebuf.sv",   "lines":"88-112",  "reason":"signal axi_wb_inflight appeared 8x"},
    {"file":"scoreboard.sv",     "lines":"245-270", "reason":"failing checker location"}
  ],
  "fingerprint_siblings": [/* past bugs with same hash, with resolutions */],
  "recent_commits": [/* git log of RTL changes in last 7 days */]
}

The composer is a ~150-line Python script. Three implementation choices matter for senior reviewers, because the naive version of each is what makes most homegrown bundle builders unreliable. RTL selection uses a pre-built symbol index (Verible, sv-parser, Surelog — pick one) scoped to the failing IP rather than blind text grep across the whole SoC; symbol matches are cached between runs and re-indexed only on RTL change. Fingerprint siblings come from a SQLite cache populated by the postmortem stage (Stage 6) and include the resolution file/line plus a short note — the LLM treats these as prior context, not as ground truth. Token budget is enforced by ranking events by causal proximity to the failure and dropping the lowest-information class first (scheduler chatter, repeated config-DB reads, idle-cycle traffic); the failure event itself and the 50 events immediately preceding it are never dropped. The script is a future post in its own right; for here, what matters is the contract the bundle obeys: token-counted, minimum-viable, structured, and reproducible from the same JSONL inputs.

Stage 4: LLM Hypothesis-Rank

Now we hand the bundle to the LLM. The single most important framing decision: ask for ranked hypotheses with falsifying experiments, not free-form analysis. Structured output is verifiable; free-form prose is not. See AI Playbook Section 5 for why this pattern beats the alternatives.

SYSTEM
You are a senior DV engineer triaging a UVM regression failure.

USER
Given the following structured failure bundle, propose the 3 most
likely root causes ranked by probability. For each cause, provide:
- a one-sentence explanation
- ONE falsifying experiment that would refute it (a plusarg, a test
  variant, a probe to add) that I can run in < 5 minutes.

Return JSON in this shape:
[
  {"cause": "...", "prob": 0.55, "falsifier": "..."},
  ...
]

Bundle:
<6,847 tokens of structured JSON>

Response shape (verbatim):

[
  {
    "cause": "D3 power gating clears axi_wb_inflight while a write is in flight",
    "prob": 0.55,
    "falsifier": "Re-run the same seed with +NO_D3_TRANSITIONS=1; if the mismatch disappears, D3 entry is the trigger"
  },
  {
    "cause": "Scoreboard predictor stale across the D0-to-D3 transition",
    "prob": 0.30,
    "falsifier": "Re-run with +SB_PREDICTOR_TRACE=1 and inspect predictor state at t=4900"
  },
  {
    "cause": "CDC race on the d3_req signal in the AXI write-buffer",
    "prob": 0.15,
    "falsifier": "Re-run with +CDC_RANDOMIZER_OFF=1; if it passes, CDC is the culprit"
  }
]

Three ranked hypotheses. Three concrete experiments, each runnable in under five minutes. No prose. We pick #1 — highest probability and easiest falsifier — and hand it to Stage 5.

Two implementation details that matter for production use. First, invoke the LLM with its structured-output mode (Anthropic tool-call interface, OpenAI JSON mode with response_format schema, or the equivalent) so the response shape is enforced by the provider rather than parsed defensively in Python. Second, pin temperature=0 and a fixed model version so the same bundle replays to the same response — without that you have a debug pipeline that is itself non-deterministic, which is the worst-of-both-worlds outcome. Cache both the prompt and the response keyed on bundle hash; senior teams that run the same failure repeatedly during a debug session do not want to pay the API cost (or wait the latency) more than once.

Stage 5: Self-Debug Loop with Smoke Test

This is where the agent earns its keep. Instead of a one-shot suggestion, the agent runs the experiment, ingests the result, and refines.

$ python -m silicondv.agent debug \
    --failure-id SCB_MISMATCH-4321-T-A8F2 \
    --bundle bundles/SCB_MISMATCH-4321-T-A8F2.json \
    --max-iterations 3

The ReAct trace, truncated to the highlights:

[Iteration 1]
REASON: Hypothesis #1 (D3 power gating clears axi_wb_inflight) ranks
        highest. Falsifier is cheap. Running.
ACT:    run_smoke(test="axi_mixed_power_test", seed=918273645,
                  plusargs="+NO_D3_TRANSITIONS=1")
OBSERVE: PASS (no mismatch). Hypothesis #1 confirmed: D3 entry triggers
         the failure.

[Iteration 2]
REASON: Now I need the precise mechanism. The context events show
        axi_wb_inflight transitions 1 -> 0 between t=4950 and t=5000,
        coincident with pwr_fsm "ENTERING_D3" -> "D3". The RTL excerpt
        for pwr_controller.sv:130-155 shows the power controller
        asserts pwr_async_rst on D3 entry. Hypothesis: pwr_async_rst is
        clearing axi_wb_inflight even when a write is in flight.
ACT:    query_rtl(pattern="axi_wb_inflight", context_lines=10)
OBSERVE: In axi_writebuf.sv:88-112, axi_wb_inflight is reset by
         pwr_async_rst directly with no gating on the inflight
         transaction.

[Iteration 3]
REASON: Root cause located. The power controller resets the AXI write
        buffer state machine without checking for in-flight
        transactions. The fix is to gate pwr_async_rst assertion on
        !axi_inflight.
PROPOSE_FIX: pwr_controller.sv:138, add condition:
             pwr_async_rst <= req_d3 && !axi_inflight;
DONE in 3 iterations / 4 min 12 sec.

Two smoke runs, one RTL query, three iterations, four minutes wall-clock. The agent did not find the bug. It narrowed a 200,000-line search space down to one signal, one file, and a small line window. The engineer (you) confirmed the exact line, made the fix, and moved on.

Two implementation notes for the senior reader. The query_rtl tool above uses a regex for illustration; the production version queries an SV abstract syntax tree (Verible, sv-parser, or Surelog) because real RTL has too many edge cases — combinational vs non-blocking assignment, macros, generate blocks, parameter expansion — for text matching to be reliable across teams. The run_smoke tool runs the simulator inside a sandbox with a hard wall-clock timeout (typically 2-3x the test's nominal runtime), structured log capture, and a cleanup hook so a runaway compile or hung VCS process does not consume the regression farm; the agent calls this tool on the LLM's recommendation and the LLM cannot be trusted to bound its own compute.

Stage 6: Postmortem Feedback

The pipeline is not over when the fix lands. The last stage closes the loop: the fingerprint goes into the known-bugs database, a coverage point is added so the regression catches a regression of the same kind, and the postmortem updates the routing table. See the Debug page §Process for the Five-Whys treatment.

$ python -m silicondv.postmortem record \
    --fingerprint "SCB_MISMATCH|axi_mixed_power|D3_entry_during_axi" \
    --resolution-file pwr_controller.sv \
    --resolution-line 138 \
    --owner pwr_team

Recorded. Routing table updated: this fingerprint -> pwr_team.
Next occurrence will auto-assign without human triage.

Next time the same fingerprint appears in a regression, the bundle stage will see it in the fingerprint-siblings list and pass that resolution context to the LLM. The pipeline learns.

The Build-It-Yourself Adoption Path

You do not adopt this pipeline in one weekend. You adopt it incrementally, in five steps, with each step independently valuable.

  1. Day 1 — Install the JSON-emitting uvm_report_server. Zero changes to test code; you get a parallel JSONL alongside every run. (Structured Logging Pattern 1)
  2. Week 1 — Add fingerprinting + run header. Triage triages itself; reproducibility becomes a one-liner. (Patterns 11-12)
  3. Week 2 — Write build_bundle.py. One-page Python: load the JSONL, take a window around an error, grep RTL files for signal names that appear, output JSON. ~150 lines. The token budget is the discipline that matters; the heuristics are negotiable.
  4. Week 3 — Wire agent_debug.py to your LLM provider. Anthropic, OpenAI, your local model — the workflow is provider-agnostic. The hypothesis-rank prompt is the durable artifact.
  5. Week 4 — Add run_smoke.py as a tool. A subprocess wrapper around your simulator. This is what closes the agent loop — without it, the LLM is a smart guesser; with it, the LLM is an iterating debugger.

Each step is small. Each step pays off immediately. By the end of the month you have a pipeline that takes a regression failure from uh oh to fixed and regression-test added in minutes — not because you bought magic, but because you built a contract-bound sequence of small, well-typed transformations.

Costs the steps above hide. The JSON-emitting uvm_report_server typically adds 5-15% to simulation wall-clock depending on log density; if your tests are perf-sensitive, the right answer is an asynchronous flush from an in-memory queue, not synchronous $fwrite on the critical path. The ring buffer holds ~1,500 events per run; on a regression farm with 10K parallel jobs that is a few GB of host RAM you did not previously budget. The bundle composer adds ~30-60 s per failure to triage time, paid once per fingerprint and cached. The LLM cost per debug session runs roughly $0.05-$0.50 at frontier-model pricing (model and bundle size dependent); budget this explicitly per regression and cache prompt+response by bundle hash so re-debug is free. None of these costs are fatal; all of them are surprises if you do not plan for them.

The next post in this series walks through one failure end-to-end — the actual prompts, the actual transcripts, the actual numbers, including the cases where the agent points you in the wrong direction. The pipeline is the system. The case study is the proof.