The Full Pipeline: From Structured Logs to AI Pair Debug in One Workflow

Every team has structured logs. Some teams have failure fingerprinting. A few teams experiment with LLM-assisted debug. Almost nobody has these three composed into a single pipeline. This post is what that composition looks like when you build it — six stages, each independently valuable from prior posts on this blog, fused into one workflow that takes a regression failure from uh oh to fixed and regression-test added in minutes instead of hours.

Read this as the integration layer between the Structured Logging for UVM post and the AI for Design Verification: A Practitioner Playbook. The patterns are not new. The composition is. Each section below names the standalone post you can read for the deeper treatment of that stage; the goal here is to show how they snap together.

The Pipeline at a Glance

Six stages. Each stage hands a small, well-typed artifact to the next. None of the boundaries are negotiable — that is what makes the pipeline composable.

flowchart TD
    A[SV testbench
structured events] --> B[uvm_report_server
+ ring buffer] B -- ERROR fires --> C[Flush: 1,500 pre-failure events] C --> D[Compute fingerprint hash
+ append to JSONL] D --> E[Python: build_bundle.py
token-budgeted context] E --> F[LLM: hypothesis-rank
top 3 causes + falsifiers] F --> G[Python: run_smoke.py
self-debug loop] G -- converges --> H[Fix + new coverage point
+ regression test] H --> I[Postmortem: fingerprint →
known-bugs DB] I -.feedback.-> A

Six artifacts, in order: a stream of structured events → a ring-buffer dump → a fingerprint → a token-budgeted bundle → a ranked-hypothesis response → a converged ReAct trace ending in a fix. Each artifact has a defined shape. Each transition is a function. Nothing is “just describe the failure to the LLM and see what happens.”

Stage 1: The Ring Buffer Fires

Your custom uvm_report_server buffers UVM_INFO events silently. On UVM_ERROR or UVM_FATAL, it flushes the last 1,500 events to a JSONL sidecar — the “context window” of the failure. See Structured Logging Pattern 9 for the full implementation.

What lands on disk:

// regression.jsonl (last 8 lines before failure)
{"ts":4800,"sev":"INFO","id":"PWR","comp":"env.pwr","msg":"entered D3","pwr_state":"D3","voltage_rail":"0.7V","freq_mhz":400}
{"ts":4850,"sev":"INFO","id":"AXI_TX","comp":"env.axi.drv","msg":"queued write","addr":"0x100","awid":3,"burst":"INCR","awlen":4,"txn_id":"4321-T-A8F2"}
{"ts":4900,"sev":"INFO","id":"AXI_RX","comp":"env.axi.mon","msg":"first beat observed","addr":"0x100","txn_id":"4321-T-A8F2"}
{"ts":4950,"sev":"INFO","id":"DUT_STATE","comp":"env.probe","msg":"snapshot","axi_wb_inflight":1,"pwr_fsm":"ENTERING_D3"}
{"ts":5000,"sev":"ERROR","id":"SCB_MISMATCH","comp":"env.axi.sb","msg":"data mismatch","addr":"0x100","exp":"0xDE","got":"0xAD","fingerprint":"SCB_MISMATCH|axi_mixed_power|D3_entry_during_axi"}

This is now your ground truth instead of human prose. Every downstream stage operates on this stream.

Stage 2: Fingerprint Collapses Buckets

Every ERROR event carries a stable fingerprint hash — failing assertion ID, last sequence kind, DUT state signature. Across a regression with 10,000 failures, fingerprints collapse the noise. See Structured Logging Pattern 12.

$ jq -r 'select(.sev=="ERROR") | .fingerprint' regression/*.jsonl \
    | sort | uniq -c | sort -rn
  4127 SCB_MISMATCH|axi_mixed_power|D3_entry_during_axi
  2891 PCIE_TLP_TIMEOUT|pcie_mem_rd_seq|L1_substate
  1450 AXI_BURST_ERROR|axi_burst_seq|gen3_x4
   823 RAL_PREDICT_FAIL|reg_access_seq|reset_release
   ...

10,234 raw failures collapse to roughly a dozen buckets. We pick the top one — SCB_MISMATCH|axi_mixed_power|D3_entry_during_axi — and feed its failure_id into Stage 3.

Stage 3: Python Pulls the Bundle

This is the first place the pipeline does something not covered by the previous posts. Python composes a context bundle — the minimum-viable, token-budgeted package the LLM will see — from the JSONL stream, the RTL, and the regression history. The bundle anatomy below is sketched; the full treatment (six-part structure, RTL excerpt selection, DUT_STATE events, multi-IP scoping, MCP-based context exposure) lives in Context Engineering for DV §9.

$ python -m silicondv.bundle build \
    --failure-id SCB_MISMATCH-4321-T-A8F2 \
    --max-tokens 8000

Bundle composed: 6,847 tokens
  - failure event: 1 event
  - context window: 187 events (ring buffer dump)
  - RTL excerpts: 3 files (1,420 tokens)
  - fingerprint siblings: 5 past failures with same fingerprint
  - recent commits: 4 RTL commits in last 7 days
Written to: bundles/SCB_MISMATCH-4321-T-A8F2.json

The bundle is one JSON object with five keys:

{
  "failure": {"ts":5000, "id":"SCB_MISMATCH", "addr":"0x100", ...},
  "context_events": [/* 187 pre-failure events, ordered */],
  "rtl_excerpts": [
    {"file":"pwr_controller.sv", "lines":"130-155", "reason":"signal pwr_fsm appeared 12x in context"},
    {"file":"axi_writebuf.sv",   "lines":"88-112",  "reason":"signal axi_wb_inflight appeared 8x"},
    {"file":"scoreboard.sv",     "lines":"245-270", "reason":"failing checker location"}
  ],
  "fingerprint_siblings": [/* past bugs with same hash, with resolutions */],
  "recent_commits": [/* git log of RTL changes in last 7 days */]
}

The composer is a ~150-line Python script. Three implementation choices matter for senior reviewers, because the naive version of each is what makes most homegrown bundle builders unreliable. RTL selection uses a pre-built symbol index (Verible, sv-parser, Surelog — pick one) scoped to the failing IP rather than blind text grep across the whole SoC; symbol matches are cached between runs and re-indexed only on RTL change. Fingerprint siblings come from a SQLite cache populated by the postmortem stage (Stage 6) and include the resolution file/line plus a short note — the LLM treats these as prior context, not as ground truth. Token budget is enforced by ranking events by causal proximity to the failure and dropping the lowest-information class first (scheduler chatter, repeated config-DB reads, idle-cycle traffic); the failure event itself and the 50 events immediately preceding it are never dropped. The script is a future post in its own right; for here, what matters is the contract the bundle obeys: token-counted, minimum-viable, structured, and reproducible from the same JSONL inputs.

Stage 4: LLM Hypothesis-Rank

Now we hand the bundle to the LLM. The single most important framing decision: ask for ranked hypotheses with falsifying experiments, not free-form analysis. Structured output is verifiable; free-form prose is not. See AI Playbook Section 5 for why this pattern beats the alternatives.

SYSTEM
You are a senior DV engineer triaging a UVM regression failure.

USER
Given the following structured failure bundle, propose the 3 most
likely root causes ranked by probability. For each cause, provide:
- a one-sentence explanation
- ONE falsifying experiment that would refute it (a plusarg, a test
  variant, a probe to add) that I can run in < 5 minutes.

Return JSON in this shape:
[
  {"cause": "...", "prob": 0.55, "falsifier": "..."},
  ...
]

Bundle:
<6,847 tokens of structured JSON>

Response shape (verbatim):

[
  {
    "cause": "D3 power gating clears axi_wb_inflight while a write is in flight",
    "prob": 0.55,
    "falsifier": "Re-run the same seed with +NO_D3_TRANSITIONS=1; if the mismatch disappears, D3 entry is the trigger"
  },
  {
    "cause": "Scoreboard predictor stale across the D0-to-D3 transition",
    "prob": 0.30,
    "falsifier": "Re-run with +SB_PREDICTOR_TRACE=1 and inspect predictor state at t=4900"
  },
  {
    "cause": "CDC race on the d3_req signal in the AXI write-buffer",
    "prob": 0.15,
    "falsifier": "Re-run with +CDC_RANDOMIZER_OFF=1; if it passes, CDC is the culprit"
  }
]

Three ranked hypotheses. Three concrete experiments, each runnable in under five minutes. No prose. We pick #1 — highest probability and easiest falsifier — and hand it to Stage 5.

Two implementation details that matter for production use. First, invoke the LLM with its structured-output mode (Anthropic tool-call interface, OpenAI JSON mode with response_format schema, or the equivalent) so the response shape is enforced by the provider rather than parsed defensively in Python. Second, pin temperature=0 and a fixed model version so the same bundle replays to the same response — without that you have a debug pipeline that is itself non-deterministic, which is the worst-of-both-worlds outcome. Cache both the prompt and the response keyed on bundle hash; senior teams that run the same failure repeatedly during a debug session do not want to pay the API cost (or wait the latency) more than once.

Stage 5: Self-Debug Loop with Smoke Test

This is where the agent earns its keep. Instead of a one-shot suggestion, the agent runs the experiment, ingests the result, and refines.

$ python -m silicondv.agent debug \
    --failure-id SCB_MISMATCH-4321-T-A8F2 \
    --bundle bundles/SCB_MISMATCH-4321-T-A8F2.json \
    --max-iterations 3

The ReAct trace, truncated to the highlights:

[Iteration 1]
REASON: Hypothesis #1 (D3 power gating clears axi_wb_inflight) ranks
        highest. Falsifier is cheap. Running.
ACT:    run_smoke(test="axi_mixed_power_test", seed=918273645,
                  plusargs="+NO_D3_TRANSITIONS=1")
OBSERVE: PASS (no mismatch). Hypothesis #1 confirmed: D3 entry triggers
         the failure.

[Iteration 2]
REASON: Now I need the precise mechanism. The context events show
        axi_wb_inflight transitions 1 -> 0 between t=4950 and t=5000,
        coincident with pwr_fsm "ENTERING_D3" -> "D3". The RTL excerpt
        for pwr_controller.sv:130-155 shows the power controller
        asserts pwr_async_rst on D3 entry. Hypothesis: pwr_async_rst is
        clearing axi_wb_inflight even when a write is in flight.
ACT:    query_rtl(pattern="axi_wb_inflight", context_lines=10)
OBSERVE: In axi_writebuf.sv:88-112, axi_wb_inflight is reset by
         pwr_async_rst directly with no gating on the inflight
         transaction.

[Iteration 3]
REASON: Root cause located. The power controller resets the AXI write
        buffer state machine without checking for in-flight
        transactions. The fix is to gate pwr_async_rst assertion on
        !axi_inflight.
PROPOSE_FIX: pwr_controller.sv:138, add condition:
             pwr_async_rst <= req_d3 && !axi_inflight;
DONE in 3 iterations / 4 min 12 sec.

Two smoke runs, one RTL query, three iterations, four minutes wall-clock. The agent did not find the bug. It narrowed a 200,000-line search space down to one signal, one file, and a small line window. The engineer (you) confirmed the exact line, made the fix, and moved on.

Two implementation notes for the senior reader. The query_rtl tool above uses a regex for illustration; the production version queries an SV abstract syntax tree (Verible, sv-parser, or Surelog) because real RTL has too many edge cases — combinational vs non-blocking assignment, macros, generate blocks, parameter expansion — for text matching to be reliable across teams. The run_smoke tool runs the simulator inside a sandbox with a hard wall-clock timeout (typically 2-3x the test's nominal runtime), structured log capture, and a cleanup hook so a runaway compile or hung VCS process does not consume the regression farm; the agent calls this tool on the LLM's recommendation and the LLM cannot be trusted to bound its own compute.

Stage 6: Postmortem Feedback

The pipeline is not over when the fix lands. The last stage closes the loop: the fingerprint goes into the known-bugs database, a coverage point is added so the regression catches a regression of the same kind, and the postmortem updates the routing table. See the Debug page §Process for the Five-Whys treatment.

$ python -m silicondv.postmortem record \
    --fingerprint "SCB_MISMATCH|axi_mixed_power|D3_entry_during_axi" \
    --resolution-file pwr_controller.sv \
    --resolution-line 138 \
    --owner pwr_team

Recorded. Routing table updated: this fingerprint -> pwr_team.
Next occurrence will auto-assign without human triage.

Next time the same fingerprint appears in a regression, the bundle stage will see it in the fingerprint-siblings list and pass that resolution context to the LLM. The pipeline learns.

The Build-It-Yourself Adoption Path

You do not adopt this pipeline in one weekend. You adopt it incrementally, in five steps, with each step independently valuable.

  1. Day 1 — Install the JSON-emitting uvm_report_server. Zero changes to test code; you get a parallel JSONL alongside every run. (Structured Logging Pattern 1)
  2. Week 1 — Add fingerprinting + run header. Triage triages itself; reproducibility becomes a one-liner. (Patterns 11-12)
  3. Week 2 — Write build_bundle.py. One-page Python: load the JSONL, take a window around an error, grep RTL files for signal names that appear, output JSON. ~150 lines. The token budget is the discipline that matters; the heuristics are negotiable.
  4. Week 3 — Wire agent_debug.py to your LLM provider. Anthropic, OpenAI, your local model — the workflow is provider-agnostic. The hypothesis-rank prompt is the durable artifact.
  5. Week 4 — Add run_smoke.py as a tool. A subprocess wrapper around your simulator. This is what closes the agent loop — without it, the LLM is a smart guesser; with it, the LLM is an iterating debugger.

Each step is small. Each step pays off immediately. By the end of the month you have a pipeline that takes a regression failure from uh oh to fixed and regression-test added in minutes — not because you bought magic, but because you built a contract-bound sequence of small, well-typed transformations.

Costs the steps above hide. The JSON-emitting uvm_report_server typically adds 5-15% to simulation wall-clock depending on log density; if your tests are perf-sensitive, the right answer is an asynchronous flush from an in-memory queue, not synchronous $fwrite on the critical path. The ring buffer holds ~1,500 events per run; on a regression farm with 10K parallel jobs that is a few GB of host RAM you did not previously budget. The bundle composer adds ~30-60 s per failure to triage time, paid once per fingerprint and cached. The LLM cost per debug session runs roughly $0.05-$0.50 at frontier-model pricing (model and bundle size dependent); budget this explicitly per regression and cache prompt+response by bundle hash so re-debug is free. None of these costs are fatal; all of them are surprises if you do not plan for them.

The next post in this series walks through one failure end-to-end — the actual prompts, the actual transcripts, the actual numbers, including the cases where the agent points you in the wrong direction. The pipeline is the system. The case study is the proof.

Structured Logging for UVM: 12 Patterns That Make Your Testbench Queryable

It is 2 AM. The overnight regression finished. 10,234 fails. You open the first log—80 MB of prose. You open the second—same shape, slightly different text. The third looks suspiciously like the first but you cannot be sure. Your job tomorrow morning: decide which of these failures are duplicates and which are real, all from human-readable strings designed to be skimmed, not queried.

With typical UVM logging, this is an eight-hour archaeology project. With structured logging, it is a single jq command that finishes before your coffee cools.

This post is the discipline that produces the second outcome. We cover 12 patterns that take any UVM testbench from prose-style logging to fully queryable. For each pattern you get a before/after code comparison, sample output, the concrete advantage you unlock, and the best-practice note that keeps you out of trouble. By the end you will have a complete drop-in uvm_report_server, a UUID-threading recipe for sequence items, a context-window ring buffer that gives you UVM_DEBUG-level forensics at UVM_LOW-level cost, and a fingerprinting strategy that collapses 10K fails into a dozen buckets in seconds.

The Problem: Logs as Prose

Compare two ways to record the same event. They cost the testbench the same to emit. They cost you very different amounts to use.

Before — prose:

`uvm_error("DRV", $sformatf("AXI burst error at addr 0x%0h, id=%0d, seq=%s",
                              addr, id, current_seq.get_name()))

What lands in the log:

UVM_ERROR @ 1234567 ns [DRV] AXI burst error at addr 0x100, id=3, seq=axi_burst_seq

Readable. Useful in a single failing run. Catastrophic at regression scale. When 10,000 failures contain variations of this line—different addresses, different IDs, different sequence names—there is no programmatic way to ask "which of these are the same bug?" except regex over prose. And the moment somebody on your team adds a tenth field with a slightly different format string, the regex breaks.

After — structured:

{"ts":1234567,"sev":"ERROR","id":"AXI_BURST_ERROR",
 "addr":"0x100","axi_id":3,"seq":"axi_burst_seq",
 "test_id":"axi_smoke_001","seed":42,"txn_id":"uuid-abc..."}

Same information. Same emission cost. Completely different downstream capabilities. jq '.addr == "0x100"' works. jq 'group_by(.id)' works. SQLite ingestion works. Sentry-style fingerprinting works. Loading into Elastic, Splunk, or a vector DB for similar-bug retrieval works. Feeding the failing slice to an LLM works.

The shift is small in code, transformative in capability. The remaining 12 patterns are extensions of this single move.

flowchart LR
    A[UVM Components] -- uvm_info / uvm_error --> B[uvm_report_server]
    B --> C[Human log
unchanged] B --> D[JSONL sidecar
structured] D --> E[jq / SQLite / Sentry] D --> F[LLM debug bundles]

Pattern 1-2: The JSON Foundation

UVM's uvm_report_server is the single emission point for every log line. Override it once, and every `uvm_info / `uvm_error in your testbench gains structure for free.

Pattern 1: JSON sidecar via custom report server. Emit two streams — the human log unchanged for engineers, a parallel JSONL file for machines.

Pattern 2: Strongly-typed log records as SV classes. Do not reach for $sformatf in components — build a typed record and let the report server serialize it. Compile-time schema enforcement is worth a lot of grief saved later.

Here is a complete drop-in implementation. Compile this into your TB and a JSONL file appears alongside every run:

// File: silicondv_log_server.sv
class silicondv_log_record extends uvm_object;
  `uvm_object_utils(silicondv_log_record)

  // Required schema fields
  string   event_id;      // e.g., "AXI_BURST_ERROR"
  string   txn_id;        // UUID threaded across components
  string   seq_name;
  longint  sim_time_ns;
  string   severity;
  string   message;

  // Optional structured fields
  string   protocol_fields[string];  // {"addr":"0x100","len":"4"}
  string   tags[$];                  // ["protocol-error","recovery"]

  function new(string name = "log_record");
    super.new(name);
  endfunction
endclass

class silicondv_report_server extends uvm_default_report_server;
  protected int unsigned m_json_fd;
  protected static string m_run_uuid;

  function new(string name = "silicondv_report_server");
    super.new(name);
    m_json_fd = $fopen("regression.jsonl", "w");
    if (m_run_uuid == "")
      m_run_uuid = $sformatf("%0t-%0d", $time, $urandom());
  endfunction

  // Override the central emission hook
  virtual function string compose_report_message(
      uvm_report_message report_message,
      string report_object_name = "");
    string human_line, json_line;
    human_line = super.compose_report_message(report_message,
                                              report_object_name);

    json_line = $sformatf(
      "{\"run\":\"%s\",\"ts\":%0d,\"sev\":\"%s\",\"id\":\"%s\",\"comp\":\"%s\",\"msg\":\"%s\"}",
      m_run_uuid,
      $time,
      severity_to_str(report_message.get_severity()),
      report_message.get_id(),
      report_object_name,
      escape_json(report_message.get_message()));

    $fwrite(m_json_fd, "%s\n", json_line);
    return human_line;  // unchanged for terminal/log file
  endfunction

  protected function string severity_to_str(uvm_severity sev);
    case (sev)
      UVM_INFO:    return "INFO";
      UVM_WARNING: return "WARN";
      UVM_ERROR:   return "ERROR";
      UVM_FATAL:   return "FATAL";
      default:     return "UNKNOWN";
    endcase
  endfunction

  protected function string escape_json(string s);
    // Minimal escaping; extend for tabs, newlines, control chars as needed
    string r = s;
    foreach (r[i])
      if (r[i] == "\"")
        r = {r.substr(0,i-1), "\\\"", r.substr(i+1,r.len()-1)};
    return r;
  endfunction
endclass

// Install in top
module tb_top;
  initial begin
    silicondv_report_server srv = new("report_server");
    uvm_report_server::set_server(srv);
    run_test();
  end
endmodule

Run a test, and now you get two artifacts side by side. The human .log looks exactly the way it always has — your colleagues can scroll it, search it, and feel at home. Alongside it sits regression.jsonl:

{"run":"4321-9921","ts":1234567,"sev":"INFO","id":"DRV","comp":"env.axi.drv","msg":"sent write 0x100"}
{"run":"4321-9921","ts":1234580,"sev":"INFO","id":"MON","comp":"env.axi.mon","msg":"observed write 0x100"}
{"run":"4321-9921","ts":1234600,"sev":"ERROR","id":"AXI_BURST_ERROR","comp":"env.axi.sb","msg":"data mismatch at 0x100"}

From this point on, every UVM message in your testbench is emitted as both human prose and a structured JSON line. The next 10 patterns are extensions of this base.

Advantages you unlock immediately:

  • Zero changes to test code. Every existing `uvm_info / `uvm_error in your codebase gains structure for free. No refactor required.
  • Human log preserved. Reviewers, scrollers, and grep-ers see no change. Adoption friction is near zero.
  • One emission point. Schema changes happen in one file, not 200 components.
  • Tool-ready out of the box. JSONL is the lingua franca of every observability stack — jq, Splunk, Elastic, Loki, BigQuery, SQLite, OpenTelemetry, vector DBs — all consume it natively.
Best practice: Keep the JSON emission cost cheap. Avoid $sformatf heroics inside compose_report_message — it runs on every message. If you need rich fields, build them in the component using the typed log record (Pattern 2) and pass them through; do not reparse the human message string on the way out.

Pattern 3-5: Identity Threading

A single failing transaction touches a sequence, a driver, the DUT, a monitor, and a scoreboard. By default each component logs independently, with no way to correlate. The fix is a transaction UUID threaded through every component that handles the transaction.

Pattern 3: UUID per uvm_sequence_item. Generate it once, in pre_randomize(), and carry it with the transaction.

class axi_txn extends uvm_sequence_item;
  rand bit [31:0] addr;
  rand bit [7:0]  data[];
  string txn_id;          // <-- threaded UUID
  string parent_txn_id;   // parent burst, parent packet, etc.

  function void pre_randomize();
    super.pre_randomize();
    if (txn_id == "")
      txn_id = $sformatf("%0t-%0d", $time, $urandom());
  endfunction

  `uvm_object_utils_begin(axi_txn)
    `uvm_field_int(addr, UVM_ALL_ON)
    `uvm_field_string(txn_id, UVM_ALL_ON)
    `uvm_field_string(parent_txn_id, UVM_ALL_ON)
  `uvm_object_utils_end
endclass

Pattern 4: Parent-child linking. When a high-level transaction expands into low-level beats (AXI burst → individual transfers, PCIe TLP → byte-level packets), copy the parent txn_id into the child's parent_txn_id. Now you can reconstruct any hierarchical operation with a single query.

Pattern 5: Sequence stack. When an error fires inside seq_a which was started by seq_b which was started by virtual sequence vseq_c, the log should carry that stack. UVM's get_sequencer().get_arbitration_sequence_q() can supply the active queue; emit it as a structured seq_stack array.

Before — correlating by address: An AXI write to 0x100 fails the scoreboard. You grep four log files looking for 0x100:

$ grep "0x100" axi_drv.log axi_mon.log scb.log seq.log
axi_drv.log: [1234567] sent write 0x100 awid=3
axi_mon.log: [1234580] observed write 0x100 awid=3
scb.log:     [1234600] mismatch at 0x100 (exp=0xDE got=0xAD)
seq.log:     [1230000] axi_burst_seq starting at 0x100
seq.log:     [1245000] axi_mixed_seq starting at 0x100  # different txn!

Which 0x100 matches which? Two sequences both touched that address. The greps return five lines but they describe two different transactions. You start cross-referencing timestamps by hand.

After — correlating by UUID:

$ jq -c 'select(.txn_id=="4321-9921-T-A8F2")' regression.jsonl
{"ts":1234567,"comp":"env.axi.drv","msg":"sent write","txn_id":"4321-9921-T-A8F2","seq_stack":["vseq_top","axi_burst_seq"]}
{"ts":1234580,"comp":"env.axi.mon","msg":"observed write","txn_id":"4321-9921-T-A8F2"}
{"ts":1234600,"comp":"env.axi.sb","msg":"mismatch","txn_id":"4321-9921-T-A8F2","exp":"0xDE","got":"0xAD"}

One query. Three lines. The complete lifecycle of the failing transaction, ordered by sim time, with the calling sequence stack attached. No cross-referencing, no false positives from another sequence touching the same address.

Advantages:

  • One query reconstructs any transaction. No more grep-by-address with false positives from address aliasing or burst overlap.
  • Hierarchical operations are inspectable. An AXI burst expanding into 16 beats? Filter on parent_txn_id and see all 16, ordered.
  • Stack-trace equivalent for verification. The seq_stack answers "what was the test doing here?" instantly — no more reading 500 cycles back to find a sequence start.
  • Stable across re-runs. Generate UUIDs deterministically from seed + transaction count and the same failure has the same UUID on replay — useful for diffing two regressions.
Best practice: Generate UUIDs that are locally unique, not globally unique. $sformatf("%0t-%0d", $time, $urandom()) beats a true RFC-4122 UUID for DV use — shorter to print, deterministic given the seed, and you do not need to coordinate across machines. The only requirement is uniqueness within a single regression batch.

Pattern 6-8: Domain Context

Generic logs are not enough for real DUTs. A bug that only happens in D3 power state at Gen3 link speeds on clock domain B needs that context attached to every relevant event — not buried in prose.

Pattern 6: Protocol-aware fields. Extend the base log record with protocol-specific structured fields. Do not dump them into a free-text message:

class axi_log_record extends silicondv_log_record;
  bit [31:0] addr;
  bit [7:0]  awid;
  bit [3:0]  awlen;
  string     burst_type;  // FIXED / INCR / WRAP

  // Serialize as structured JSON fields, not embedded strings
  virtual function string to_json_fields();
    return $sformatf(
      ",\"addr\":\"0x%0h\",\"awid\":%0d,\"awlen\":%0d,\"burst\":\"%s\"",
      addr, awid, awlen, burst_type);
  endfunction
endclass

Pattern 7: Power state. Every log line should carry the current power state (D0/D1/D2/D3, voltage rail, current frequency). Most teams put this in uvm_info prose; promoting it to a structured field makes "all failures in D3 at 800 MHz" a one-line query.

Pattern 8: Clock domain ID. Multi-clock designs benefit enormously from a clk_domain field. Errors that cluster by clock domain reveal CDC bugs that are otherwise invisible in flat logs.

Implementation pattern: maintain a singleton log_context object that components update on state transitions. The report server reads the current context on every emission and merges it into the JSON.

Before — the "only fails on Tuesdays" bug: A scoreboard mismatch appears in 11 of 800 nightly tests. Always different addresses, always different sequences. You spend two days hunting for a common factor. Eventually someone notices that the failing logs all happen to contain the substring "D3" buried in a comment 200 lines earlier — they were all in low-power state.

UVM_ERROR @ 5000 ns [SCB] mismatch at 0x100 (exp=DE got=AD)
UVM_INFO  @ 4800 ns [PWR] entered D3 state, rail=0.7V freq=400MHz   # 200 lines back
...

The context is in the log—but in prose, on a different line, in a different file. Discoverable only by accident.

After — context attached at the source:

{"ts":5000,"sev":"ERROR","id":"SCB_MISMATCH","addr":"0x100","exp":"0xDE","got":"0xAD",
 "pwr_state":"D3","voltage_rail":"0.7V","freq_mhz":400,
 "clk_domain":"clk_b","link_mode":"gen3_x4"}

Now the discovery is a single query:

$ jq -c 'select(.sev=="ERROR") | {id, pwr_state, freq_mhz}' regression.jsonl \
    | sort | uniq -c | sort -rn
   2891 {"id":"SCB_MISMATCH","pwr_state":"D3","freq_mhz":400}
     42 {"id":"SCB_MISMATCH","pwr_state":"D0","freq_mhz":800}
     11 {"id":"SCB_MISMATCH","pwr_state":"D3","freq_mhz":800}

One look and you know: this bug is overwhelmingly a low-power-state issue at 400 MHz, not a generic scoreboard issue. Two days of hunting compressed into 30 seconds of querying.

Advantages:

  • Mode-specific bugs surface fast. "Always in D3," "only on clock domain B," "only at Gen3 x4" become one-line queries instead of forensic exercises.
  • CDC bugs become visible. Errors clustered by clk_domain reveal cross-domain races that are otherwise lost in a flat timeline.
  • Context is preserved at emission time. Even if a downstream tool strips other fields, the structured columns survive ingestion into databases and dashboards.
  • Cross-protocol triage works. The same query strategy applies whether the failing IF is AXI, PCIe, USB, or your custom protocol — only the protocol-specific fields differ.
Best practice: Update the log_context singleton from events, not from state polling. When the DUT enters D3, an event-driven update writes once; a polled context that asks the DUT on every log line is a measurable simulation slowdown. The same principle applies to clock domain, link mode, and reset state — update on transitions, read on emission.

Pattern 9: The Context-Window Ring Buffer

This is the single highest-value pattern in this post. The problem: you want UVM_DEBUG-level detail at the point of failure, but you cannot afford UVM_DEBUG verbosity for the entire run. Disk fills, simulation slows, signal-to-noise collapses.

The solution, borrowed from SWE observability (Sentry breadcrumbs, Honeycomb context capture): keep the last N events in a ring buffer, suppress them from disk output, and on UVM_ERROR or UVM_FATAL, flush the entire buffer.

flowchart LR
    A[INFO event] --> B{Severity?}
    B -- INFO --> C[Push to ring
DEPTH=1500] B -- ERROR/FATAL --> D[Flush buffer
+ emit current] C --> E[(suppressed)] D --> F[JSONL output]
class log_ring_buffer #(int DEPTH = 1500);
  protected string m_buf[$];

  function void push(string json_line);
    m_buf.push_back(json_line);
    if (m_buf.size() > DEPTH) void'(m_buf.pop_front());
  endfunction

  function void flush(int fd);
    foreach (m_buf[i]) $fwrite(fd, "%s\n", m_buf[i]);
    m_buf.delete();
  endfunction
endclass

class silicondv_report_server extends uvm_default_report_server;
  log_ring_buffer #(.DEPTH(1500)) m_ring = new();
  int m_json_fd;

  virtual function string compose_report_message(
      uvm_report_message rm, string obj = "");
    string json_line = build_json(rm, obj);

    if (rm.get_severity() inside {UVM_ERROR, UVM_FATAL}) begin
      // Flush 1500 events of pre-failure context, then current
      m_ring.flush(m_json_fd);
      $fwrite(m_json_fd, "%s\n", json_line);
    end else if (rm.get_severity() == UVM_INFO) begin
      // Buffer INFO; emit higher severities directly
      m_ring.push(json_line);
    end

    return super.compose_report_message(rm, obj);
  endfunction
endclass

Before — the verbosity dilemma: Run at UVM_LOW and the JSONL stays small, but when an error fires you have no idea what happened in the 500 cycles leading up to it. Run at UVM_DEBUG and you get the context, but the JSONL grows to 4 GB per test, the simulator slows by 30-50%, and your regression farm runs out of disk by morning. So you compromise — reproduce the failure with high verbosity. Except half the time the failure does not reproduce.

After — ring buffer flushes on error:

// regression.jsonl during the run: nothing
// (1500 INFO events sitting silently in the ring buffer)

// Then at t=5000 ns, UVM_ERROR fires. Buffer flushes:
{"ts":3502,"sev":"INFO","comp":"env.axi.drv","msg":"sent write 0xF8","txn_id":"...A8E0"}
{"ts":3580,"sev":"INFO","comp":"env.axi.mon","msg":"observed write 0xF8","txn_id":"...A8E0"}
{"ts":3600,"sev":"INFO","comp":"env.pwr","msg":"entered D3","pwr_state":"D3"}
{"ts":4000,"sev":"INFO","comp":"env.axi.drv","msg":"sent write 0x100","txn_id":"...A8F2"}
... 1495 more events ...
{"ts":5000,"sev":"ERROR","comp":"env.axi.sb","msg":"mismatch","id":"SCB_MISMATCH"}

You get a complete pre-failure timeline: every transaction, every state change, every coverage hit, every assertion that fired in the 1,500 events leading up to the error. The signal that mattered is right there. The 9,500,000 events earlier in the run that did not matter are not.

Advantages:

  • UVM_DEBUG-level forensics at UVM_LOW-level disk cost. A clean run produces a tiny JSONL. A failing run produces a focused forensic dump.
  • Reproduces are not required. The context you would have wanted from a high-verbosity replay is already captured the first time.
  • Simulator stays fast. No $fwrite per INFO event under steady state — the ring buffer is in-memory until something fails.
  • Works on emulation. The same pattern survives the move to Veloce or ZeBu, where the disk-cost calculus is even more brutal.
Best practice: Tune the depth to your design. 1,500 events is a sensible default for IP-level testbenches; SoC-level may want 5,000-10,000 to span the failure window. Measure: take a few real failures from your last regression, count how many events backward you actually scrolled to find the cause, then set DEPTH to 2x that. Going too deep wastes memory; going too shallow truncates the cause.

Pattern 10: Coverage Integration

Functional coverage is the verification team's ground truth for "what did we test." But the linkage between coverage and logs is usually broken — you know a covergroup hit bin 47, but not which test, seed, or transaction triggered it.

The fix is small: auto-emit a structured log on every covergroup sample, with the bin name and value attached. Hook into the same code path that calls sample():

// In your covergroup wrapper class
function void sample_with_log(input axi_txn txn);
  cg.sample();
  log_event(.event_id("COVERAGE_HIT"),
            .txn_id(txn.txn_id),
            .fields($sformatf(
              "\"cg\":\"axi_cg\",\"sample_addr\":\"0x%0h\"",
              txn.addr)));
endfunction

Before: The coverage report says bin axi_cg::burst_len_wrap_16 is hit. Great—but which test? Which seed? Which transaction inside that test? The answer requires re-running with verbose coverage tracing, which often runs slower than the original test and produces gigabytes of output you then have to comb through.

After: Every bin hit emits a structured event with the triggering transaction's txn_id. Now the inverse is one query:

$ jq -r 'select(.id=="COVERAGE_HIT" and .cg=="axi_cg") \
         | "\(.test_id) seed=\(.seed) bin=\(.bin_name)"' regression.jsonl \
    | sort -u
axi_smoke_001 seed=42  bin=burst_len_wrap_16
axi_mixed_007 seed=918 bin=burst_len_wrap_16
axi_stress_03 seed=33  bin=burst_len_wrap_16

Three tests hit the bin. You know which seeds. You know which transactions (filter by txn_id for the full lifecycle). Sign-off readiness becomes inspectable.

Advantages:

  • Coverage holes get test attribution. Combine the coverage database with the coverage-event log to find "bins hit by only one test" or "bins hit only at high seed counts" — both flags for fragile coverage.
  • Cross-reference with failures. Did the failing transactions also hit a coverage bin? jq 'select(.txn_id=="...")' shows the bins it touched on its way to failure.
  • Regression diffing. Coverage events from yesterday's run vs. today's surface unintended coverage drops from a code change.
Best practice: Do not emit a coverage event for every sample if your covergroup samples every cycle. Emit only on first hit of a new bin — track which bins have already been logged in a thread-local set. Otherwise the JSONL grows by an order of magnitude with no added signal.

Pattern 11: The Run Header

Agans's Rule 2 — Make It Fail — requires that any bug be reproducible. The run header is the structured artifact that makes reproduction trivial: a single log event emitted at start_of_simulation_phase containing everything needed to reproduce.

function void start_of_simulation_phase(uvm_phase phase);
  super.start_of_simulation_phase(phase);
  log_event(.event_id("RUN_HEADER"), .severity("INFO"),
            .fields($sformatf(
              "\"seed\":%0d,\"tool_version\":\"%s\",\"rtl_hash\":\"%s\",\"tb_hash\":\"%s\",\"plusargs\":\"%s\"",
              $get_initial_random_seed(),
              `SIMULATOR_VERSION,
              `RTL_GIT_HASH,
              `TB_GIT_HASH,
              get_plusargs_dump())));
endfunction

Before — "can you reproduce that?": Someone files a bug at 5 PM Friday. To reproduce on Monday you need the seed, the RTL git hash, the testbench git hash, the simulator version, and the exact +plusargs that were active. These live, respectively, in the regression metadata JSON, the build manifest, a separate build manifest, the simulator's banner output (already rotated out of disk), and a Makefile that has been modified since. You spend half a day reconstructing the run.

After — one event captures everything:

{"ts":0,"sev":"INFO","id":"RUN_HEADER",
 "seed":918273645,
 "tool_version":"VCS 2026.03-1",
 "rtl_hash":"a4c91e8",
 "tb_hash":"7b3d2f1",
 "plusargs":"+UVM_TESTNAME=axi_mixed_test +AXI_GEN3=1 +PWR_D3_TIME=200"}

To reproduce: one jq against the failing run's JSONL, one git checkout, one vcs invocation.

$ jq -r 'select(.id=="RUN_HEADER")' failing.jsonl | \
    jq -r '"git checkout \(.rtl_hash) && simv +UVM_TESTNAME=\(.plusargs) -ntb_random_seed=\(.seed)"'
git checkout a4c91e8 && simv +UVM_TESTNAME=+UVM_TESTNAME=axi_mixed_test... -ntb_random_seed=918273645

This single event is the entire "make it fail" replay artifact. Combined with seed-stamped log lines (Pattern 15 in the appendix), reproducing any failure becomes a one-liner.

Advantages:

  • Reproduction stops being archaeology. Five disparate sources collapse to one structured event at start_of_simulation.
  • Bisection is mechanical. Combined with Pattern 12 (fingerprints), git bisect can run a smoke test against each RTL commit until the failure shape changes.
  • Bug reports get better. When the RUN_HEADER lands in the bug ticket, the reporter cannot accidentally omit a critical detail.
  • Compliance traceability. For projects that need audit trails (automotive, medical, aerospace), every failure carries provable provenance.
Best practice: Emit the RUN_HEADER at start_of_simulation_phase, not run_phase. If your test fails in build_phase (config DB miss, factory override mismatch), the header was already written and the failure is reproducible. Late emission defeats the purpose.

Pattern 12: Failure Fingerprints

Sentry, Rollbar, and Bugsnag pioneered the idea that errors should be grouped automatically by a stable hash. The same principle works for verification: emit a fingerprint with every error, derived from the failing assertion ID, the last sequence kind, and a DUT state signature. Errors that share a fingerprint are duplicates.

function string compute_fingerprint(uvm_report_message rm, string last_seq_kind);
  // Stable across re-runs; identifies the failure shape, not the run instance
  return $sformatf("%s|%s|%s",
                   rm.get_id(),
                   last_seq_kind,
                   get_dut_state_signature());
endfunction

With this field in every error JSON, your triage script collapses 10,234 failures into roughly a dozen unique fingerprints. The same algorithm that takes hours of human deduplication takes seconds:

jq -r 'select(.sev=="ERROR") | .fingerprint' regression/*.jsonl \
  | sort | uniq -c | sort -rn

Sample output:

  4127 AXI_BURST_ERROR|axi_burst_seq|state_idle_to_active
  2891 PCIE_TLP_TIMEOUT|pcie_mem_rd_seq|state_l0
  1450 SCB_MISMATCH|axi_mixed_seq|state_active
   823 RAL_PREDICT_FAIL|reg_access_seq|state_idle
   ...

You went from "10,234 failures, unknown structure" to "12 buckets, ranked by frequency" in one command.

Before — the spreadsheet of doom: A senior engineer opens 50 log files in tabs, copies the error string out of each, pastes them into a spreadsheet, sorts, manually groups similar strings, and assigns each cluster to a sub-team. Eight hours of work per regression. Done by the person whose time is most valuable. Repeated every morning. The team budgets a half-day of "triage tax" into every sprint.

After: The pipeline above runs at the end of the regression. The next morning, you open a dashboard showing 12 fingerprint buckets ranked by frequency, each with the owning sub-team auto-assigned by matching the fingerprint prefix against a routing table. Triage tax: minutes.

Advantages:

  • Triage scales sub-linearly. A 10x regression produces only marginally more buckets — the bug count is bounded by real defects, not by run count.
  • Hot bugs are obvious. The bucket with 4,127 occurrences gets investigated first; the bucket with 11 occurrences either is a real corner case or gets waived.
  • Regression-to-regression diffs work. If a fingerprint appears today that did not appear yesterday, you have a true new regression — not just a re-occurrence of a known bug.
  • Auto-routing. Map fingerprint prefixes to owners and JIRA/GitHub issue assignment becomes mechanical.
  • Dedup across machines. Five regression nodes each finding the same bug independently no longer triples your apparent failure count.
Best practice — what makes a good fingerprint: The three components matter. Failing assertion ID identifies what broke. Last sequence kind identifies what was running. DUT state signature (a hash of key FSM states + mode + power) identifies under what conditions. Skip any of the three and dissimilar bugs collide; add more and identical bugs split. Tune the state signature on real data — start with FSM state alone, then add power state, then add link mode, watching the bucket count converge to a sensible number.

16 More Patterns (for future posts)

This post covered the dozen that move the needle most. The catalog continues:

  1. Severity-stratified files — split errors.jsonl, warnings.jsonl, info.jsonl for fast CI gating.
  2. Versioned schema field — backward compatibility when the log shape evolves.
  3. Seed in every log line — not just the run header, so any cut log is still reproducible.
  4. Pre/post state-diff snapshots — emit before/after register state on key events.
  5. SVA-triggered structured events — bind SVA fires to log emissions with triggering signal values.
  6. Latency histograms per transaction type — emit summary events every N transactions.
  7. Stall reason codes — structured cause tracking for every backpressure event.
  8. Hierarchical log namespacing — soc.cpu.l1cache.miss for SoC bring-up.
  9. Per-IP verbosity control — debug one IP at HIGH while others stay LOW.
  10. Plusarg log filters — +log_filter=axi,!pcie for surgical noise reduction.
  11. NDJSON streaming output — one event per line; ideal for tail -f + jq pipelines.
  12. SQLite-loadable schema — run SQL across an entire regression's logs.
  13. OpenTelemetry export — feed protocol traces into Jaeger or Tempo.
  14. Verdi / DVE TCL commands embedded — one-click jump from log line to waveform.
  15. DPI-C bridges to Python logging — live tail of testbench events in your SWE tooling.
  16. Pre-built LLM debug bundles — token-budget-aware extraction for GenAI debug.

Each of these will get its own deep-dive post — see the Debug page for the full methodology map.

Putting It Together

You can adopt structured logging incrementally. The order that maximizes payoff per hour of effort:

  1. Day 1: Install the JSON-emitting uvm_report_server (Pattern 1). Zero changes to test code; you get a parallel JSONL file alongside every run.
  2. Day 2: Add UUID stamping to your sequence item base class (Pattern 3). All transactions are now correlatable.
  3. Week 1: Emit a run header (Pattern 11). Every failure is now reproducible from a single log event.
  4. Week 2: Add the context-window ring buffer (Pattern 9). The single highest-value debug capability in this post.
  5. Week 4: Implement failure fingerprinting (Pattern 12). Regression triage stops being a person-week and starts being a script.

Each step compounds with the others. By the end of the month, your testbench logs are queryable, your transactions are traceable, your failures are deduplicated, and your debug context is automatic. Every later technique on the Debug page — Delta Debugging, GenAI context bundles, AI-assisted similar-bug retrieval — builds on this foundation.

Structured logging is not a tool. It is the discipline that makes every other debug technique possible.