AI Pair Debug, Worked: An AXI Burst Mismatch in D3 Power State

This is a worked case study. A failing test, walked end-to-end through the pipeline from The Full Pipeline post. The prompts are real verbatim. The LLM responses are illustrative but representative of what a frontier model produces given this bundle shape. The numbers at the end are real measurements from running this workflow on similar bugs in a working DV environment.

Read this as evidence the workflow described in the previous post actually works — and an honest accounting of where it does not. If you want the patterns abstracted, read the previous post first. If you want to see what the pattern feels like in practice, this is that.

Note: the specific bug below is a constructed scenario — an AXI burst mismatch during a D3 power state entry. It was chosen because it is realistic, has clear before/after stages, exercises three different subsystems (AXI, power, scoreboard), and avoids any IP or customer disclosure. The pipeline patterns and the LLM behavior are real.

The Setup

Test: axi_mixed_power_test, seed 918273645. A randomized sequence of AXI write bursts interleaved with periodic power-state transitions between D0 (active) and D3 (deep sleep). Lasts roughly 1.2 million cycles.

Failure: One scoreboard mismatch at simulation time 5,000 ns. Expected data 0xDE, observed 0xAD, at address 0x100. Single failure in an overnight regression of 800 tests — rare enough to be hard to dismiss, frequent enough that letting it slip is not an option.

What the structured log shows (the last 8 events before the error, from the ring-buffer dump):

{"ts":4800,"sev":"INFO","id":"PWR_STATE","comp":"env.pwr","msg":"entered D3","pwr_state":"D3","voltage_rail":"0.7V","freq_mhz":400}
{"ts":4850,"sev":"INFO","id":"AXI_TX","comp":"env.axi.drv","msg":"queued write","addr":"0x100","awid":3,"burst":"INCR","awlen":4,"data_first":"0xDE","txn_id":"4321-T-A8F2"}
{"ts":4900,"sev":"INFO","id":"AXI_RX","comp":"env.axi.mon","msg":"first beat observed on bus","addr":"0x100","awid":3,"txn_id":"4321-T-A8F2"}
{"ts":4920,"sev":"INFO","id":"DUT_STATE","comp":"env.probe","msg":"snapshot","axi_wb_inflight":1,"pwr_fsm":"ENTERING_D3","axi_wb_addr":"0x100"}
{"ts":4960,"sev":"INFO","id":"DUT_STATE","comp":"env.probe","msg":"snapshot","axi_wb_inflight":0,"pwr_fsm":"D3","axi_wb_addr":"0x000"}
{"ts":4980,"sev":"INFO","id":"AXI_RX","comp":"env.axi.mon","msg":"completion observed","addr":"0x100","bresp":"OKAY","txn_id":"4321-T-A8F2"}
{"ts":5000,"sev":"ERROR","id":"SCB_MISMATCH","comp":"env.axi.sb","msg":"data mismatch","addr":"0x100","exp":"0xDE","got":"0xAD","txn_id":"4321-T-A8F2","fingerprint":"SCB_MISMATCH|axi_mixed_power|D3_entry_during_axi"}

Traditional triage estimate: The senior engineer assigned this would open the FSDB, hunt for axi_wb_inflight transitions around 4920 ns, correlate with the power FSM trace, walk the AXI write-buffer module, and probably find this in 3-4 hours. We are going to do it in under 30 minutes.

First Attempt: The Naive Paste

Before we use the pipeline, let us show what happens with the “just paste the log” approach — the move every team starts with and that the AI Playbook warns against.

USER (verbatim)
Here is a failing UVM regression log from an AXI test. What is wrong?

[paste of last 2,000 lines of human log, ~38,000 tokens]

The response (excerpted; the full thing is ~600 tokens of mostly hedging):

ASSISTANT
Looking at the log, several things stand out:

1. There is an AXI write that the scoreboard reports as a mismatch.
   This could be:
   - A timing issue on the AXI bus
   - A race condition in the driver or monitor
   - A scoreboard predictor issue
   - An issue with the burst encoding

2. The error occurs near a power state transition. This could be
   unrelated, but worth checking.

3. I would recommend enabling more verbose AXI tracing and checking
   for protocol violations...

To narrow this down further, could you share:
- The relevant AXI driver code
- The scoreboard implementation
- The exact RTL module being tested
- The simulator version

Generic. Hedging. Asks for four follow-up artifacts the engineer already has in front of them. Mentions the power transition as “could be unrelated” — missing that it is in fact the trigger. This is what an unstructured paste produces from any frontier model: the context-window-degradation effect (arxiv 2509.21361) plus the lack of a focused prompt frame combine to produce a response that is not wrong, but is not useful.

Lesson: the model is not going to compensate for your context discipline. You have to bring it.

Second Attempt: With a Context Bundle

Now we use the pipeline. First, build the bundle (the bundle's structure, ordering rules, and selection heuristics are deep-treated in Context Engineering for DV §9):

$ python -m silicondv.bundle build \
    --failure-id SCB_MISMATCH-4321-T-A8F2 \
    --max-tokens 8000

Bundle composed: 6,847 tokens
  - failure event: 1 event
  - context window: 187 events (ring buffer dump)
  - RTL excerpts: 3 files (1,420 tokens)
      * pwr_controller.sv:130-155 (signal pwr_fsm hit 12x)
      * axi_writebuf.sv:88-112    (signal axi_wb_inflight hit 8x)
      * scoreboard.sv:245-270     (failing checker location)
  - fingerprint siblings: 0 past failures with this fingerprint
  - recent commits: 4 RTL commits in last 7 days
Written to: bundles/SCB_MISMATCH-4321-T-A8F2.json

6,847 tokens — an order of magnitude smaller than the naive paste. Now the prompt:

SYSTEM
You are a senior DV engineer triaging a UVM regression failure.

USER
Analyze this structured failure bundle and identify the most likely
root cause. Be specific about which signal, file, and timing window
you suspect.

[bundle, 6,847 tokens]

Response (excerpted):

ASSISTANT
The failure is highly correlated with a D3 power state entry. From the
context events:

- At t=4800, the system enters D3 (voltage_rail=0.7V, freq=400MHz).
- At t=4850, an AXI write to 0x100 is queued.
- At t=4920, the DUT_STATE probe shows axi_wb_inflight=1, pwr_fsm=
  ENTERING_D3.
- At t=4960, axi_wb_inflight is unexpectedly 0 while pwr_fsm=D3 —
  the write is still in flight on the bus but the internal flag has
  been cleared.
- At t=4980, the AXI completion observes on the bus.
- At t=5000, the scoreboard reports a mismatch.

My primary hypothesis: the power controller is asynchronously clearing
the AXI write-buffer state during D3 entry, even though a write is
mid-flight on the bus. This would explain the corrupted data.

Looking at the RTL excerpt for pwr_controller.sv:130-155, the
pwr_async_rst signal is asserted on D3 entry. I would investigate
whether this signal has an unconditional path to the AXI write
buffer's state machine reset.

Recommendation: examine pwr_controller.sv around the pwr_async_rst
assertion logic, and axi_writebuf.sv around the inflight-flag reset
in its always_ff block.

A genuinely useful response. Names the right subsystems (power controller, AXI write buffer). Identifies the relevant signal (pwr_async_rst). Points at the right files within 4-10 lines of the actual fix location. Still: it is one hypothesis, asserted as the primary, with no falsifier. A senior engineer wants verification before touching RTL. We tighten the frame.

Iteration 1: Hypothesis-Rank

We re-prompt with the structured-output discipline from the AI Playbook:

USER
Thank you. Reframe your analysis as the 3 most likely root causes
ranked by probability. For each cause, provide ONE falsifying
experiment I can run in <5 minutes (a plusarg, test variant, or
probe).

Return JSON in this shape:
[
  {"cause": "...", "prob": 0.XX, "falsifier": "..."},
  ...
]

Response:

[
  {
    "cause": "Power controller's pwr_async_rst clears axi_wb_inflight without checking for in-flight transactions",
    "prob": 0.55,
    "falsifier": "Re-run the same seed with +NO_D3_TRANSITIONS=1; if the mismatch disappears, D3 entry is the trigger"
  },
  {
    "cause": "Scoreboard predictor is stale across the D0-to-D3 transition and reports a stale expected value",
    "prob": 0.30,
    "falsifier": "Re-run with +SB_PREDICTOR_TRACE=1 and inspect predictor state at t=4900; if the expected value differs from 0xDE, the predictor is the issue"
  },
  {
    "cause": "CDC race on d3_req signal between the power clock domain and the AXI clock domain",
    "prob": 0.15,
    "falsifier": "Re-run with +CDC_RANDOMIZER_OFF=1; if it passes, CDC is the culprit"
  }
]

Three ranked hypotheses. Three concrete experiments, each runnable in under five minutes. Hypothesis #1 is highest probability and has the cheapest falsifier (one plusarg flip). We start there.

Iteration 2: Self-Debug Loop

We invoke the agent to run the experiments and converge.

$ python -m silicondv.agent debug \
    --failure-id SCB_MISMATCH-4321-T-A8F2 \
    --bundle bundles/SCB_MISMATCH-4321-T-A8F2.json \
    --max-iterations 3

The ReAct trace (excerpted to the highlights):

[Iteration 1]
REASON  Hypothesis #1 ranks highest. The falsifier is one plusarg.
        Running.
ACT     run_smoke(test="axi_mixed_power_test", seed=918273645,
                  plusargs="+NO_D3_TRANSITIONS=1")
OBSERVE PASS in 47 seconds. No mismatch. D3 entry confirmed as the
        trigger. Hypothesis #1 confirmed; #2 and #3 not yet falsified
        but lower priority.

[Iteration 2]
REASON  D3 entry is the trigger. Now I need the precise mechanism.
        The DUT_STATE events at t=4920 and t=4960 show axi_wb_inflight
        transitions from 1 to 0 between them, coincident with
        pwr_fsm transitioning ENTERING_D3 to D3. From the RTL excerpt
        for pwr_controller.sv:130-155, pwr_async_rst is asserted on
        the D3 transition. I need to see how axi_wb_inflight is reset
        in the write-buffer module.
ACT     query_rtl(pattern="axi_wb_inflight\\s*<=", context_lines=12)
OBSERVE In axi_writebuf.sv around line 92:
            always_ff @(posedge clk or posedge pwr_async_rst) begin
              if (pwr_async_rst) axi_wb_inflight <= 1'b0;
              else if (axi_aw_handshake) axi_wb_inflight <= 1'b1;
              else if (axi_b_handshake)  axi_wb_inflight <= 1'b0;
            end
        The pwr_async_rst is in the asynchronous reset list with no
        gating on in-flight state. Confirmed root cause.

[Iteration 3]
REASON  Root cause located. The power controller's reset clears the
        write-buffer state without regard for in-flight transactions.
        The data the AXI monitor observes (0xAD) is whatever was on
        the bus after the reset corrupted the buffer; the scoreboard's
        expected value (0xDE) is what was queued before the reset.
PROPOSE_FIX  In pwr_controller.sv near line 138, gate the assertion
             of pwr_async_rst on absence of in-flight AXI activity:
             pwr_async_rst <= req_d3 && !axi_inflight_top;
             Alternatively, in axi_writebuf.sv, remove pwr_async_rst
             from the always_ff sensitivity list and replace with a
             synchronous, conditional reset that respects the
             handshake. The first fix is smaller; the second is more
             defensively designed.

DONE in 3 iterations / 4 min 12 sec.
Proposed fix saved to: fixes/SCB_MISMATCH-4321-T-A8F2.diff

Two smoke runs, one RTL query, three iterations, four minutes. The agent did not invent the bug; it narrowed a 200,000-line search space down to one signal, one file, and a 4-line window. It even proposed two alternative fixes with a brief design judgment.

The Fix and Validation

The engineer (you) opens pwr_controller.sv and looks at the proposed line. The agent said line 138; the actual fix turns out to be at line 134 — the agent was off by a few lines but in the right neighborhood. The fix is a one-line change:

// Before
pwr_async_rst <= req_d3;

// After  
pwr_async_rst <= req_d3 && !axi_inflight_top;

Smoke test passes. Full regression re-runs in 30 minutes — passes. A new coverage point is added so the regression catches a regression of the same kind:

// In the power coverage group
cover (d3_entry_req && axi_inflight_top) {
  bins suppressed_during_axi = {1};
  bins fired_during_axi      = {0};  // should never increment
}

And the postmortem is recorded:

$ python -m silicondv.postmortem record \
    --fingerprint "SCB_MISMATCH|axi_mixed_power|D3_entry_during_axi" \
    --resolution-file pwr_controller.sv \
    --resolution-line 134 \
    --owner pwr_team \
    --notes "pwr_async_rst gated on !axi_inflight_top"

Recorded. Routing table updated.

Next time this fingerprint appears, the bundle stage will surface the resolution in the fingerprint-siblings list, and the LLM will incorporate it as prior context.

What Worked, What Did Not

An honest accounting matters more than a victory lap. Here is what landed and what did not.

What worked

  • The bundle stage rescued the LLM. The same model produced near-useless prose from a raw log paste and a sharp, well-aimed hypothesis from a 6,847-token bundle. The context-engineering principle is real and measurable.
  • Hypothesis-rank produced a falsifiable starting point. Asking for ranked hypotheses with falsifiers (rather than analysis) forced structured output that we could act on without thinking too hard.
  • The self-debug loop converged in 3 iterations. Two smoke runs were enough to confirm the hypothesis and locate the root cause to a 4-line RTL window.
  • The RTL excerpt selector got lucky and useful. The heuristic that picks RTL by signal-name hit count put pwr_controller.sv and axi_writebuf.sv in front of the LLM — without that, the model would have asked us to paste the code.
  • Reproducibility was built in. Every LLM call used temperature=0 and a pinned model version, so the same bundle replays to the same response; the full agent transcript was logged with bundle hash, prompt, response, and tool-call sequence so the session is auditable after the fact. For regulated flows (automotive, medical, aerospace) this is non-negotiable; for everyone else it is what makes the session debuggable when the agent goes off the rails.

What did not

  • The naive prompt was a waste. Pasting the raw log produced a generic, hedging response. Without the bundle, the pipeline does not function. Anyone trying to skip steps 3-5 will conclude the LLM “does not work for DV.”
  • 2 of 3 hypotheses were wrong. The rank is a starting point, not a verdict. If hypothesis #1 had been wrong, we would have spent another smoke run on #2. Treat the ranks as priorities for the engineer, not as truth from the model.
  • The proposed RTL line number was off by ~4 lines. Frontier models are not reliable at exact line-number precision. The agent gave us the neighborhood (within a 25-line excerpt); we found the exact line. This is the right division of labor — do not expect more.
  • The model proposed two fixes; both were viable but only one was idiomatic for this team. Architectural judgment on which fix to take is still ours. The agent narrowed the choice; we made it.

Time Comparison

Measured end-to-end times for this category of bug, from one engineer running similar workflows on similar bugs over a regression season.

StepTraditionalThis workflow
Identify failure cluster30 min (read logs by hand)10 sec (jq + uniq -c)
Build mental model of failure45 min (waveform navigation)30 sec (build_bundle)
Hypothesize root cause30 min (think + sketch)2 min (LLM hypothesis-rank)
Verify hypothesis30 min (manual probe + re-run)6 min (2 smoke runs at ~3 min each)
Locate exact RTL line30 min (grep + read)2 min (agent's RTL query + human verification)
Apply fix + validate30 min10 min
Total triage time~3 hours~25 minutes

Not in the comparison: the one-time pipeline setup cost (estimated 1-2 weeks for a working JSON-emitting uvm_report_server, bundle composer, and agent harness). The full regression re-run (~30 minutes, but it happens in either workflow). Code review on the fix (~15 minutes, again the same). The LLM API cost for this debug session (~$0.10-$0.30 at frontier-model pricing — trivial in absolute terms but worth budgeting at scale). The ~5-15% simulation overhead the JSON-emitting report server adds during regression, and the few GB of host RAM consumed by ring buffers across a parallel regression farm; both are real and both have mitigation patterns (async flush, smaller ring depth for non-debug runs) that you implement once and forget.

The 25-minute number is door-to-door triage time after the pipeline is in place. Your mileage will vary: some bugs are easier (the agent solves them in 5 minutes; some are harder (the agent points you in the wrong direction and you fall back to traditional methods). The honest average across a season is roughly 10x faster than the traditional baseline — less for easy bugs, more for those where the LLM context-engineering rewards your investment.

The workflow, not the model, is the durable thing. Swap Claude for GPT-5 for Gemini for your local model — the bundle, the hypothesis-rank, the self-debug loop, the smoke gate are model-agnostic. Pick your provider. Commit to the workflow.