Error Fingerprinting for UVM Regressions: Bucketing 10,000 Failures the Way Sentry and Windows Do

The overnight regression finished with 10,234 failures. The structured-logging post ended with a one-line jq that collapses them into a dozen buckets by hashing three fields. That line works, and it is also where most teams stop. Then, two weeks later, the buckets start lying: a scoreboard message that hides five different bugs, an address in a message that splits one bug into four hundred buckets, a hang that lands in whichever bucket the watchdog message happens to match. Nobody trusts the dashboard, and the senior engineer is back to opening logs in tabs.

The software world hit exactly this wall twenty-five years ago and wrote down what it learned. Microsoft's Windows Error Reporting has bucketed crash reports from a billion machines since 1999. Sentry and Rollbar bucket exceptions for hundreds of thousands of applications and expose their grouping rules in public documentation. Microsoft Research published the algorithm that replaced message matching with call-stack matching. This post takes those results and rebuilds error fingerprinting for UVM regressions on top of them: what a fingerprint must be orthogonal to, what the "stack trace" of a testbench failure actually is, how to normalise a message without destroying its signal, and how to keep the rules honest as the project changes.

Two ways a fingerprint lies

The Windows Error Reporting team defined the goal in one sentence: a bucketing algorithm should maintain orthogonality, one bug per bucket and one bucket per bug. Every failure of a fingerprint is a failure of one half of that sentence.

  • Under-split: two bugs in one bucket. Your scoreboard reports SCB_MISMATCH for every data error it sees. A parity bug in the write path and a byte-lane bug in the read path both surface as SCB_MISMATCH from the same component. Bucketed on message id and component, they are one bucket with 1,450 hits, one owner, and one very confused afternoon.
  • Over-split: one bug in many buckets. The same scoreboard prints the address in the message. One bug, hit by 400 seeds at 400 addresses, becomes 400 buckets of one hit each. The dashboard says "400 distinct problems" and the team stops reading it.

The Vennsa and University of Toronto paper that called failure triage "the neglected debugging problem" drew exactly these two pictures for hardware: two distinct bugs caught by the same checker, and one bug caught by different checkers because different stimulus propagated it along different paths. Their observation was that the industry's automation, where it existed, binned "purely on the error message and the owners of the failing tests", and that this was the source of both failure modes.

WER's vocabulary for fixing this is worth adopting as-is. A heuristic that adds information to the fingerprint is expanding: it increases the bucket count so that two bugs stop sharing one. A heuristic that removes information is condensing: it decreases the bucket count so that one bug stops spanning several. Their table of client-side heuristics is mostly expanding (add the module name, the offset, the exception code, the hang wait-chain root) with a few deliberate condensing ones (replace the module and offset with an in-code assert ID, because the assert ID identifies the bug better than where it fired). Every change you make to a fingerprint rule is one of these two moves, and you should know which one you are making and why.

DV translation The equivalent of WER's assert_tags condensing heuristic is using the SVA property name or the uvm_report id instead of the message text. The equivalent of the expanding hang_wait_chain heuristic is fingerprinting a timeout by which sequence was blocked and on what, not by the watchdog's generic message.

The failure stack: what your testbench's stack trace is

Sentry's default grouping does not look at the message first. It looks at the stack trace, and it distinguishes in-app frames, your code, from library and framework frames, which it ignores for grouping. Rollbar does the same: its exception fingerprint is a hash of the exception class plus the file and method names of the frames, with line numbers dropped and framework boilerplate frames removed. The 2012 ReBucket paper from Microsoft Research went further and replaced exact matching with a similarity measure: frames near the top of the stack (the crash point) weigh more than frames near the bottom, and two stacks that match the same functions at slightly different depths still count as similar. ReBucket also strips "immune functions", library code that is trusted enough to be unlikely to hold the bug.

A UVM failure has a stack trace too. It is just not in the place a software engineer would look. The frames, from top to bottom, are:

FrameSourceWeightWhy
Assertion or checker idrm.get_id(), property nameHighestClosest to the observation of the bug, the "crash point"
Reporting componentrm.get_report_object().get_full_name(), instance indices maskedHighWhich monitor or scoreboard saw it
Phaseuvm_phase at the time of the errorMediumReset-phase failures and run-phase failures are rarely the same bug
Running sequence chainget_parent_sequence() walked to the rootMedium, decaying downwardWhat stimulus was active; the outer virtual sequence matters less than the innermost
Test name+UVM_TESTNAMELowestThe frame most scripts bucket on, and the least informative: one test hits many bugs and one bug hits many tests

Notice the weight column is the inverse of common practice. Grouping by test name is grouping by the bottom frame, the main() of the failure. ReBucket's position-dependent model says that is the frame to trust least.

The in-app rule also transfers cleanly. Components in your environment are in-app. Anything under uvm_pkg and inside a purchased VIP is a library frame: keep it for context, drop it from the fingerprint. A failure reported by uvm_test_top.env.pcie_vip.dl_layer.crc_checker is fingerprinted at the VIP boundary, env.pcie_vip, plus the VIP's own error id, because the frames inside it are not yours to fix and their internal paths change with every VIP release.

Here is the collector. It runs inside a report catcher, so it sees every error at the moment it is raised, and it emits one compact record per failing simulation, on the first error only.

class failure_stack_catcher extends uvm_report_catcher;
  `uvm_object_utils(failure_stack_catcher)

  bit             captured;
  string          library_roots[$] = '{"pcie_vip", "axi_vip"};
  // Set from phase_started() in your base test: current_phase = phase.get_name();
  static string   current_phase = "build";

  function new(string name = "failure_stack_catcher");
    super.new(name);
  endfunction

  // Mask instance indices: env.axi_agent[3].monitor -> env.axi_agent[*].monitor
  function string mask_indices(string path);
    string out = "";
    foreach (path[i]) begin
      if (path[i] == "[") begin
        out = {out, "[*]"};
        while (i < path.len() && path[i] != "]") i++;
      end else out = {out, path[i]};
    end
    return out;
  endfunction

  // First index of sub in s, or -1 (SystemVerilog has no built-in substring search)
  function int index_of(string s, string sub);
    for (int i = 0; i + sub.len() <= s.len(); i++)
      if (s.substr(i, i + sub.len() - 1) == sub) return i;
    return -1;
  endfunction

  // Cut the path at a library boundary so VIP internals do not enter the fingerprint
  function string to_in_app(string path);
    foreach (library_roots[k]) begin
      int idx = index_of(path, library_roots[k]);
      if (idx >= 0) return path.substr(0, idx + library_roots[k].len() - 1);
    end
    return path;
  endfunction

  function action_e catch();
    uvm_report_object ro;
    uvm_sequence_base seq;
    string frames[$];
    if (captured || get_severity() < UVM_ERROR) return THROW;
    captured = 1;

    ro = get_client();
    frames.push_back({"id:", get_id()});
    frames.push_back({"comp:", mask_indices(to_in_app(ro.get_full_name()))});
    frames.push_back({"phase:", current_phase});

    // Sequence chain, innermost first, stored as kinds not instances
    if ($cast(seq, ro) == 0) seq = running_sequence(ro);
    while (seq != null) begin
      frames.push_back({"seq:", seq.get_type_name()});
      seq = seq.get_parent_sequence();
    end

    log_failure_record(frames, get_message());
    return THROW;
  endfunction
endclass

Three hooks are yours. Override phase_started() in the base test with one line, failure_stack_catcher::current_phase = phase.get_name();, so the catcher knows the phase without reaching into UVM internals. running_sequence() and log_failure_record() are the other two: the first resolves the sequence currently on the sequencer that drives the reporting agent, the second appends one JSON line to the sidecar file described in the structured-logging post. The record carries the frames, the raw message, and the run header fields (seed, test, RTL hash, tool version). Everything below consumes that record.

Normalising the message without losing the signal

Rollbar publishes its message-normalisation list, and it is short: strip dates, timestamps, email addresses, IP addresses, decimal numbers, integers of two or more digits, and hex values of four or more digits, then hash what is left. The interesting part is the exception: error codes and HTTP status codes are deliberately kept, because a 404 and a 500 in the same handler are different bugs. "Strip things that look like data, keep things that look like codes" is the whole rule.

The DV version of the list:

MaskPatternKeep or stripReason
Simulation time@ 12345 ns, time=...StripChanges with every seed
Seedseed=..., +ntb_random_seedStripRun identity, not bug identity
Addresses and data0x[0-9a-f]{4,}, 'h...StripThe classic over-split source
Transaction and packet idsid=NN, seq_id, tag=StripPer-run counters
Instance indices[3] in pathsStripSame bug across ports
Small integers0 to 9 standing aloneKeepUsually a lane, a channel, a state: a code, not data
State and opcode namesIDLE, L0s, WRITE_BURSTKeepThe DV equivalent of an error code
Checker verdict wordsexpected, actual, timeout, underflowKeepDistinguish mismatch from protocol violation

The one judgment call is the two-digit rule. Rollbar strips integers of two or more digits and keeps single digits. That heuristic is right for hardware too: a lane index of 3 is a code, a burst length of 256 is data. Apply the same rule and you get 90 percent of the benefit with no per-message configuration.

import re

MASKS = [
    (r"@\s*\d+\s*[pnuf]?s\b",              "@<T>"),    # sim time
    (r"\b(seed|ntb_random_seed)\s*[=:]\s*\d+", r"\1=<SEED>"),
    (r"0x[0-9a-fA-F]{4,}",                 "<HEX>"),    # addresses, data
    (r"'h[0-9a-fA-F_]{4,}",                "<HEX>"),
    (r"\b(id|tag|seq_id|txn)\s*[=:]\s*\d+", r"\1=<ID>"),
    (r"\[\d+\]",                           "[*]"),      # instance indices
    (r"\b\d{2,}\b",                        "<N>"),      # integers of 2+ digits
]

def normalise(msg: str) -> str:
    for pat, rep in MASKS:
        msg = re.sub(pat, rep, msg)
    return re.sub(r"\s+", " ", msg).strip()

# "SCB_MISMATCH addr=0x4000_1000 lane 3 expected 0xdead actual 0xbeef @ 128340 ns"
#   -> "SCB_MISMATCH addr=<HEX> lane 3 expected <HEX> actual <HEX> @<T>"

If your logs are prose because the testbench predates structured logging, you do not have to write masks by hand. The Drain algorithm, published in 2017 and maintained as the drain3 package, mines message templates from a stream of log lines: variable positions become <*> automatically, a similarity threshold decides when a line starts a new template, and custom masks for hex and integers are applied first. Feed it your UVM_ERROR lines and the template id it returns is a usable normalised message on day one.

Label in the simulation, classify in the script

WER's most transferable design decision is that bucketing happens in two phases in two places. Labeling runs on the client, from the evidence available at the moment of the crash, and is cheap. Classifying runs on the server, once more data has arrived (symbols, memory dumps, a second report from the same bucket), and can move a report to a better bucket. Bucketing is progressive: the first answer is not the final answer.

Your simulation is the client. Your triage script is the server.

  • Level 1, in the simulation: a label from immediate evidence only. Top two frames plus the normalised message: id|component|template. It costs nothing, needs no other run, and is right often enough to route the failure.
  • Level 2, in the triage script: a classification that adds evidence the simulation could not see cheaply or at all. The last five transaction kinds from the ring buffer. The DUT state signature. Whether the failure was preceded by a reset or a power transition. Whether the same seed passed on yesterday's RTL. This is where the scoreboard's five hidden bugs get separated, because the read-path bug always has a READ_BURST in the last five kinds and the write-path bug never does.

Two records that share a Level 1 label but split at Level 2 tell you the label was under-split, and that is an expanding heuristic waiting to be promoted into the simulation-side rule. Two Level 2 buckets that a human keeps merging tell you the opposite. The levels are not just a pipeline. They are how the rule set learns.

From the literature WER's !analyze tool grew to roughly 100,000 lines implementing about 500 bucketing heuristics, at about one new heuristic a week for a decade. That number is the honest forecast for any fingerprint rule set: it will keep changing. Store the rule-set version in every record, and never compare buckets across versions without re-bucketing the old records.

Rules, overrides and merges

Sentry and Rollbar both expose the same three controls, and their syntaxes are worth copying because they were arrived at by years of users asking for the same things.

Fingerprint rules match an event and replace the fingerprint. Sentry writes them as one line each, first match wins:

error.type:DatabaseUnavailable -> system-down
error.value:"connection error: *" -> connection-error, {{ transaction }}
logger:my.package.* level:error -> error-logger, {{ logger }} title="Error from Logger {{ logger }}"

Rollbar writes them as JSON with a condition, a fingerprint template and a title, and its condition language has the operators you would expect: equality, membership, prefix and suffix, numeric comparison, regex, combined with any, all and none. Both let a rule refer back to the default fingerprint, so a rule can refine grouping instead of replacing it.

Merges are the human override for under-split rules: two issues merged so that future events land in one, with an unmerge that restores the originals. Sentry now also proposes merges from a transformer embedding of stack traces, which is where "similar enough" grouping ends up once exact matching runs out of road.

For a regression flow the same three controls fit in one YAML file next to the testbench, evaluated by the triage script in order:

version: 7
rules:
  # Condensing: every VIP-internal CRC error is one bug at the VIP boundary
  - match: { id: "PCIE_DL_CRC*", comp: "env.pcie_vip*" }
    fingerprint: "pcie-vip-crc|{{ phase }}"
    title: "PCIe VIP CRC error"

  # Expanding: the scoreboard hides read-path and write-path bugs behind one id
  - match: { id: "SCB_MISMATCH", last_kinds: "*READ_BURST*" }
    fingerprint: "{{ default }}|read-path"

  # Timeouts are grouped by what was blocked, never by the watchdog text
  - match: { id: "WATCHDOG_TIMEOUT" }
    fingerprint: "hang|{{ seq[0] }}|{{ comp }}"
    title: "Hang in {{ seq[0] }}"

merges:
  - into: "a1f9c2e0"
    from: ["7b3d0e11", "c04e9a77"]
    reason: "Same byte-enable bug seen by monitor and scoreboard, JIRA DV-418"

The evaluator is short. Conditions are shell-style globs against record fields, templates pull fields into the fingerprint string, and the hash is taken over the final string so that rules can be reordered without changing existing bucket ids.

import fnmatch, hashlib, json, re, yaml

def matches(rule, rec):
    return all(fnmatch.fnmatch(str(rec.get(k, "")), pat)
               for k, pat in rule["match"].items())

def render(template, rec, default):
    def sub(m):
        key = m.group(1).strip()
        if key == "default":
            return default
        if key.startswith("seq["):
            return rec["seq"][int(key[4:-1])] if rec.get("seq") else ""
        return str(rec.get(key, ""))
    return re.sub(r"{{\s*(.*?)\s*}}", sub, template)

def fingerprint(rec, cfg):
    default = "|".join([rec["id"], rec["comp"], rec["template"]])   # Level 1 label
    fp, title = default, rec["id"]
    for rule in cfg["rules"]:
        if matches(rule, rec):
            fp = render(rule["fingerprint"], rec, default)
            title = render(rule.get("title", title), rec, default)
            break
    digest = hashlib.sha1(fp.encode()).hexdigest()[:8]
    for m in cfg.get("merges", []):
        if digest in m["from"]:
            digest = m["into"]
    return digest, fp, title, cfg["version"]

Three properties of this design matter more than the code. Rules live in version control next to the testbench, so a rule change is reviewed like any other change. The rule-set version travels with every record. And the merge table is data, not a special case, so an override made at 2 AM survives the next rule edit.

The first error is not always the root

Every fingerprint pipeline starts by taking the first error in the log. Two situations break that assumption, and both have a software precedent.

Hangs. WER's hang heuristic does not fingerprint a hang by the fact that the UI stopped. It walks the chain of threads waiting on synchronisation objects, starting from the input thread, to find the root of the wait chain, and buckets on that. The DV equivalent: a UVM_FATAL from the watchdog is the observation, not the bug. The fingerprint should come from the watchdog's structured report of what was blocked, the sequence that never got its response, the sequencer it was waiting on, and the phase, which is why the rule above fingerprints WATCHDOG_TIMEOUT on seq[0] and comp rather than on the message. If your watchdog does not record that, that is the first thing to fix, and it is the subject of the Watchdog card on the Debug page.

Cascades. The first error in file order is the first error in simulation time only if you have one log. With several agents logging to sidecars, take the earliest by simulation time across all of them. Then prefer the frame closest to the root: an assertion failure on the interface beats a scoreboard mismatch two thousand cycles later, the same way ReBucket weights the top frame over the frames below it. If the assertion and the mismatch always occur together, that is a merge, not a rule.

The academic DV work is about pushing evidence closer to the root than any log message can get. The Vennsa paper built signatures from the excitation and propagation paths reported by a root-cause-analysis engine. Poulos and Veneris at ITC 2014 represented each failure as a feature vector over SAT-derived suspect sets and toggle-frequency windows, then clustered, reporting 89 percent binning accuracy and 47 percent fewer misplaced failures than message-based scripts. VCDiag in 2025 classifies failures from compressed VCD waveforms and names the top three suspect modules with over 94 percent accuracy. None of that is in reach for a nightly script today, but all of it slots into the same place: as additional fields on the Level 2 record, consumed by the same rules file. Build the pipeline so that a richer signature is a new column, not a new system.

Statistics as a debugging tool

The WER team's mantra was "data not decibels". The point of buckets is not the list; it is what you can compute over the list once the buckets are stable.

  • Pareto. WER found that a small number of buckets account for most reports. ReBucket's data for one product: 87 percent of buckets held 20 percent of hits, and 13 percent of buckets held 80 percent. Your regression will look the same. Fix the top three buckets and the tail will rise into view, which is the correct order of work.
  • New versus known. A fingerprint that appears today and did not appear yesterday is a regression. A fingerprint that appears every day is a known issue. Those two lists are the morning report, and they are a set difference between two files.
  • Bucket age and trend. First seen, last seen, hits per night. A bucket whose count is falling after a fix went in confirms the fix. A bucket whose count is rising after a fix went in is a different bug wearing the same fingerprint, which is an expanding-rule request.
  • One-hit wonders. WER reported about 10 percent of buckets with exactly one report. Do not waive them by count. A one-hit bucket that is new is a corner case the constraint solver found once; keep the seed and re-run it before deciding. A one-hit bucket that is old and never recurred is where waivers belong.
  • Owner routing. A rules file can carry an owner per fingerprint prefix, and that is enough to start. One DVCon paper trained a random forest on historical signature-to-owner assignments and reported near-perfect owner prediction for known signatures, weaker on unseen ones. The fingerprint history you accumulate is exactly that training set, so the routing table becomes a model when the table stops scaling.

The pipeline, end to end

flowchart LR
  A["Simulation
failure_stack_catcher"] -->|"one JSONL record per fail
Level 1 label"| B["Sidecar files
regression/*.jsonl"] B --> C["normalise()
mask time, seed, hex, ids"] C --> D["Level 2 evidence
last kinds, DUT state, phase"] D --> E["rules.yaml
fingerprint(), merges"] E --> F["Bucket table
count, first/last seen, owner"] F --> G["Diff vs yesterday
new / known / fixed"] G --> H["Routing
owner per bucket, one issue per bucket"]

The triage script joins the pieces. It reads every sidecar, normalises, computes the fingerprint under the current rule set, and writes a bucket table plus a diff against the previous night.

import glob, json, os, collections

def load_records(pattern):
    for path in glob.glob(pattern):
        with open(path) as f:
            for line in f:
                rec = json.loads(line)
                if rec.get("sev") in ("ERROR", "FATAL") and rec.get("first_error"):
                    rec["template"] = normalise(rec["msg"])
                    yield rec

def bucket(records, cfg):
    table = collections.defaultdict(lambda: {"hits": 0, "seeds": [], "tests": set()})
    for rec in records:
        digest, fp, title, ver = fingerprint(rec, cfg)
        b = table[digest]
        b.update(fp=fp, title=title, rules_version=ver)
        b["hits"] += 1
        if len(b["seeds"]) < 3:
            b["seeds"].append(rec["seed"])           # representative repro seeds
        b["tests"].add(rec["test"])
    return table

def diff(today, yesterday):
    new   = [k for k in today if k not in yesterday]
    fixed = [k for k in yesterday if k not in today]
    known = [k for k in today if k in yesterday]
    return new, known, fixed

cfg = yaml.safe_load(open("rules.yaml"))
today = bucket(load_records("regression/*.jsonl"), cfg)
yesterday = json.load(open("buckets/latest.json")) if os.path.exists("buckets/latest.json") else {}
new, known, fixed = diff(today, yesterday)

for k in sorted(today, key=lambda k: -today[k]["hits"]):
    b = today[k]
    flag = "NEW" if k in new else "   "
    print(f"{flag} {k} {b['hits']:5d}  {b['title']:32s}  seeds={b['seeds']}")

On a real nightly the output looks like this. The numbers are from a 190-failure run on an AXI subsystem with the rules file above at version 7.

NEW 3e9a71c0    84  SCB_MISMATCH read-path            seeds=[8812, 1207, 5590]
    a1f9c2e0    41  Byte-enable bug (merged)          seeds=[311, 9920, 4471]
    7702bd14    27  Hang in axi_burst_write_seq       seeds=[1002, 1003, 1044]
    c8d0e5f2    18  PCIe VIP CRC error                seeds=[42, 77, 5610]
    12ab8c93     9  AXI_RESP_DECERR lane 3            seeds=[6001, 6002, 6100]
    5f4e2a17     5  RAL_PREDICT_FAIL run              seeds=[913, 2211, 7370]
NEW 90c1d3ab     3  SVA a_wready_within_16            seeds=[4488, 4489, 4490]
    e77b0146     2  SCB_MISMATCH write-path           seeds=[1500, 1501]
NEW 0b6f9d55     1  UVM_FATAL cfg missing in agent[*] seeds=[2]

190 failures, 9 buckets, 3 new, 0 fixed since 2026-09-15

Read the top line the way WER would. A new bucket with 84 hits on the first night it appears is the regression; it gets the first engineer of the day. The merged byte-enable bucket is known and has an owner. The one-hit fatal with seed 2 is a config bug in a new agent instance, obvious from the title, five minutes to fix. Three people, three buckets, and nobody opened a log to decide that.

Measuring whether your buckets are honest

ReBucket was evaluated with an F-measure over hand-labelled crash data, and you can do the same at a scale that fits a Friday afternoon. Take fifty failures from one regression and have the people who fixed them write down the real bug id next to each. Then compute two numbers over your buckets:

  • Purity: for each bucket, the fraction of its failures that belong to its majority bug. Low purity means under-split; you need an expanding rule.
  • Inverse purity: for each real bug, the fraction of its failures that landed in its majority bucket. Low inverse purity means over-split; you need a condensing rule or a merge.

Keep the fifty labelled failures as a regression test for the rules file. Every rule change re-runs against them, and a change that raises purity while dropping inverse purity gets the same review a testbench change would. ReBucket reported an F-measure of about 0.88 on Microsoft's data with a learned similarity model; a hand-written rules file on a testbench you control should reach that on the second iteration, because you have something Microsoft never had: the ability to add a field to the record at the source.

Quick reference

IdeaWhere it comes fromIn the testbench
One bug per bucket, one bucket per bugWindows Error ReportingName every rule change as expanding or condensing
Group by stack, not message; top frames weigh moreSentry, Rollbar, ReBucketFailure stack: id, component, phase, sequence chain, test; test name last
In-app versus library framesSentry, ReBucket immune functionsCut paths at the VIP or uvm_pkg boundary
Strip data, keep codesRollbarMask time, seed, hex, ids, indices; keep state and opcode names and single digits
Template mining for prose logsDraindrain3 over UVM_ERROR lines when no sidecar exists
Label cheaply, classify with more evidenceWER two-phase bucketingLevel 1 in the report catcher, Level 2 in the triage script
Rules, titles, merges as dataSentry and Rollbar rule syntaxrules.yaml with version, in the testbench repo
Hangs bucket on the wait chainWER hang_wait_chainFingerprint timeouts on blocked sequence and component
Pareto, new versus known, one-hit wondersWER statisticsBucket diff every night; re-run new one-hit seeds before waiving
Purity and inverse purityReBucket evaluationFifty labelled failures as the rules file's regression test

Further reading

AI Coverage Closure: What Actually Closes Bins

Every claim in the vendor decks is true, and most of them are not about closing coverage. That is the puzzle this post untangles. AI tools for coverage closure are real, deployed, and in a few verified cases spectacular — one production SoC interconnect went from a 79.91% plateau to 100% functional coverage at 18x less compute. But the marquee numbers you have seen — 16x, 10x, 5x — are almost all compression results: the same coverage, reached cheaper. Closing new coverage is a different problem, the published wins at it are rarer and more conditional, and the conditions are the interesting part. This post maps the manual closure loop you already run, sorts the AI offerings into three tiers by what they actually touch, walks through the verified production results, and then spends equal time on what none of them automate — because the honest boundary of these tools is exactly where your judgment still lives.

The Last-Mile Problem

Start with the loop as every DV team actually runs it, stated with unusual candor by an NVIDIA team at DVCon: identify covergroups, code the bins, run regressions, generate reports and analyze, "modify constraints manually and rerun the regressions," and — step six — "repeat steps 4 and 5 until we hit 100%." Behind that "repeat" sits the whole machinery: merge the coverage databases (urg, imc, vcover merge), rank the tests, triage the holes, categorize each one — needs a new test, needs a constraint tweak, genuinely unreachable, or waivable — write exclusions, get them reviewed, rerun. The loop is not broken. It just has a cost curve that turns hostile precisely when you need it most.

The best production dataset ever published on that curve comes from a Samsung + Cadence paper on a mobile application processor's SoC interconnect — hundreds of masters, near a thousand slaves, hundreds of thousands of functional coverage elements (DVCon US 2025). One regression pass: 274 runs, 1,744 CPU-hours, 26.74% coverage. Twenty iterations: 5,480 runs, 33,751 CPU-hours, 94.84%. The next four iterations — 1,096 more runs, 7,408 more CPU-hours — bought 1.43 points. Do the division and the efficiency collapse is about 15x: the first 94.84% cost roughly 356 CPU-hours per coverage point; the tail cost about 5,180. After 41,158 CPU-hours, 4,478 bins were still open.

A software engineer will recognize the shape instantly: it is a feedback loop with exponentially decaying reward. Constrained-random stimulus re-hits the easy bins the way CI re-runs already-passing tests — every regression pays full price to reconfirm what the last one proved, and the probability of landing on an unhit bin shrinks as the unhit set does. Infineon quantified the redundancy on a production radar DSP in their Aurix line: closing coverage took on the order of two million tests, a thousand machines and licenses running nearly continuously for six months — and an industry-standard ranking algorithm afterward showed that about 3,000 of those tests would have sufficed to hold 100% coverage (DVCon US 2021). Over 550 tests simulated for every one that added coverage.

This is the pressure the AI tools are selling into, and the pressure is real. The 2024 Wilson Research Group study — still the latest as of this writing — has first-silicon success at 14%, the lowest in two decades of tracking; 75% of IC/ASIC projects behind schedule; design engineers spending 49% of their time on verification. Nobody needs convincing that the loop is expensive. What needs examining is which part of it each tool actually touches.

Your Regression Is 30–60x Bigger Than Its Coverage

Before any AI enters the picture, understand what plain test ranking already proves, because it is the baseline every ML claim should be measured against — and in the one independent study that did measure, the baseline nearly tied.

Ranking is the greedy algorithm your coverage tools have shipped for years: given a merged, test-associated coverage database, select the minimal set of test-seed pairs that reproduces the merged coverage. Infineon ran it across three real projects with Cadence (DVCon) and the numbers are startling: a microprocessor IP regression of 260 runs compressed to 8 runs at 100% coverage regain — 32.5x. A mixed-signal SoC's 5,124 runs compressed to 1,204, again at full regain. Your regression is 30–60x larger than its minimal coverage-proving subset, and the tooling to prove it is already in your license.

Two caveats turn this from a party trick into engineering judgment. First, ranking hits exactly the bins the original regression hit — never one more. It is compression by construction, and a compressed regression cannot close a hole. Second — and this is the finding teams skip past — the same study showed the compressed regressions changed the failure profile: one stage went from 3 failing runs in the original to 23 in the optimized set. The redundant tests were not worthless; they were soak. The coverage-proving subset and the bug-finding regression are different objects, and a tool that optimizes the first while you silently assume it preserved the second is how a "more efficient" flow ships a bug.

When the same study benchmarked Cadence's Xcelium ML against this non-ML ranking baseline, the result was deflationary and useful: "Both Xcelium ML and Ranking methods gave comparable compression & speedup factors around 3 consistently" — with ranking sometimes compressing more. The ML tool's one structural advantage: its regenerated regressions occasionally exercised genuinely new scenarios, regaining more than 100% of the original coverage (101–108% in several configurations) — something ranking cannot do by construction. Hold that thought; it is the entire difference between the next two sections.

One more piece of the baseline: hole triage has five buckets, not four. Needs a new test; needs a constraint or seed change; genuinely unreachable (route it to formal, not to a human); waivable with review; and — the bucket most flows don't have — the coverage model itself is wrong. A Samsung Memory team measured that last bucket on a production cache-managing IP: 14.1% of their coverage holes were bins that were never defined at all — the initial hand-written model covered only 53.5% of the true bin space (DVCon). Keep that number in mind when a tool promises to close your holes. Some of your holes are not holes.

Three Tiers of "AI Coverage" — Read the Fine Print

Every commercial "AI coverage" offering does one of three things, and the tier determines what the tool can possibly deliver. Sorting the market this way is the single most useful filter you can apply to a vendor deck.

Tier A — selects or ranks existing tests. Same stimulus pool, fewer cycles. AMD's SNUG result with Synopsys VSO.ai — 1.5–16x fewer tests to the same coverage across four designs — lives here, as do Renesas's 2.2x/3.6x Xcelium ML compressions, VSO.ai's regression-ROI ordering, and the change-based smoke-suite selection Intel presented in 2026. A Tier A tool, by construction, cannot close a coverage hole. It can only make the coverage you already reach cheaper — which is genuinely valuable, and is where nearly every marquee number comes from.

Tier B — steers the constrained-random distribution. The tool reaches into the randomization kernel or constraint solver and re-weights what stimulus gets generated, from the same testbench and the same constraints. Xcelium ML's accelerated-closure mode, Cadence's Verisium SimAI, and VSO.ai's in-simulator coverage-directed solving live here. Tier B can hit bins that plain random rarely reaches — this is where the real closure results live — but it cannot reach anything your constraints exclude. It explores your legal space more cleverly; it does not enlarge it.

Tier C — authors new verification artifacts. Tests, sequences, properties that did not exist before. As of mid-2026 no dedicated coverage-closure product from the big three is in this tier. What is here: the research wave (agentic property generation, LLM testbench synthesis) and the vendors' new agentic layers — Cadence's ChipStack, Synopsys's AgentEngineer, Siemens's Questa One Agentic Toolkit, all announced within a single month in early 2026, all early-access, all pitched at bring-up and productivity rather than last-mile closure.

The reader's rule that falls out: for every number in a vendor deck, ask did coverage go up, or did the same coverage get cheaper? Renesas's split is also worth carrying with you — 3.6x compression on a derivative design versus 2.2x on the original, because the ML had regression history to learn from. These tools eat your data, and a team with deep regression archives has a moat a fresh project does not. As one DAC 2026 wrap-up put it, the EDA companies build the tools, but the training data belongs to the companies building chips.

What Actually Closed New Coverage

Now the wins that survive fact-checking — each one primary-sourced, each with its condition attached.

The Samsung + Cadence interconnect result is the strongest closure number in public. On the second covergroup category — the one where the traditional flow plateaued at 79.91% after 41,158 CPU-hours — the SimAI-guided flow reached 100% in 2,261 CPU-hours: 18.2x less compute and 20.09 points more coverage, from the same testbench. On the first category, where the baseline had already clawed to 96.27%, the gain was 4.92x. Read those two numbers together: the advantage was largest where the tail was worst. The tail is where these methods earn their keep, not where they fail. ("Up to 18x" is doing real work in the abstract, though — the two categories are the whole spread.)

A second Samsung team reported the same shape on a production camera-interface IP: 134 cross covergroups, 14,616 bins; the ML-driven flow reached 100% in about 75,000 tests while the baseline regression "fails to exceed 70% coverage rate, even after more than 100,000 runs."

Infineon's novelty-driven selection is the best-documented Tier A result: an autoencoder ranks candidate tests by reconstruction error — novelty, in effect anomaly detection pointed at your own stimulus — and simulation order follows novelty. 60% fewer tests to reach 99.5% coverage, still 40% fewer at 99.95%, projecting the six-month closure campaign to under three months. Two honesty flags the paper itself carries: the time saving is a projection from an offline replay of 85,470 already-generated tests, not a measured deployment; and the savings decay from 60% to 40% exactly in the final half-percent, where the hard bins live.

NVIDIA's VSO.ai deployment — designs exceeding 100 million coverage targets — reported 33% more functional coverage in the same number of runs, alongside a 5x regression-suite reduction. Their published flow combines test grading, formal unreachability analysis, and VSO.ai; the often-quoted 17%-more-coverage-at-3.5x-compression figure belongs to that combined flow, not to the AI tool alone. Their lesson learned is also on the record: the tool was first tried late in a project, and the vendor now recommends deploying at early milestones, while stimulus is immature — late-stage deployment underperforms.

On the research side, one result is worth more than all the benchmark tables: LLM4DV, the Cambridge/Imperial/lowRISC benchmark for LLM-driven stimulus generation (FCCM 2025). In its 2023 version, GPT-3.5 managed 5.61% coverage on the hardest DUT, an Ibex CPU. In the current version, on identical scaffolding, Claude 3.5 Sonnet reaches 100% on that same CPU — against a constrained-random baseline of 15.31%. Across the benchmark the per-model spread runs from 7.93% to 98.84% on the same design with the same framework. The framework was never the binding constraint. The model was. Whatever you concluded about LLM stimulus generation from a 2023-era evaluation, the conclusion has a shelf life measured in model releases.

The Catch: A Human Still Writes the Targets

Here is the pattern connecting every win above, and it is the thesis of this post: the automation is in reaching specified targets efficiently, not in discovering what to target. The coverage model and the scenario list remain human artifacts, and every failure mode in this section is a way of forgetting that.

The Samsung + Cadence paper says it outright: the DV team supplies the target scenario specification, and "if the target specification is incomplete… the proposed approach may not hit the bins even though all the test scenarios of the provided specification are satisfied." Deriving the specification automatically is listed as future work. The 18x result is a solver being steered brilliantly toward targets a human enumerated.

Nokia and MathWorks hit the boundary from the other side (DVCon Europe 2023). Their autoencoder test selection cut tests-to-closure by up to 43% — but the coverage goal was capped at 67%, "the maximum possible for the current test randomization constraint configuration." No selection strategy, however intelligent, could touch the remaining third, because the constraints excluded it. Selection cannot fix a constraint problem. (The paper's abstract says "2x speedup"; its own results section reports that wall-clock regression time was not improved — doubled, in the worst case — because each iteration restarted the simulation environment. Cite this paper carefully.)

The Samsung missing-bins result gives the model-side version: 14.1% of holes were bins nobody wrote. An AI aimed at your holes is aimed at the defined bin set. The gap between the bins you wrote and the bins you should have written is invisible to every tool in this post — as the Verilab crew put it in the best paper ever written on coverage quality, "if something is missing from the model, it does not appear as a coverage hole, it's simply invisible." Their companion warning belongs on a wall: functional coverage modeling "is essentially a software task. As such, models will likely have bugs… Given the trust we put in functional coverage results for tapeout decisions, this is an oddly overlooked requirement." Mark Litterick's Lies, Damned Lies, and Coverage names the failure taxonomy — deception, omission, fabrication — and the reason coverage bugs survive: "if you make a mistake in stimulus or checks, there is a good chance you will kill the regression; if you mess up coverage there is no comeback." Goodhart's law — when a measure becomes a target it ceases to be a good measure — comes from economics via the software-testing literature, but DV built its own sharper version first.

Now put an AI in that loop and watch the metric detach from the goal. Infineon's agentic formal-coverage work (arXiv:2603.03147) is admirably honest about what happened: LLM agents read Jasper coverage reports, characterized the uncovered RTL, and generated new SVA properties, lifting formal coverage by roughly 10–20% across five designs. And: "in some cases… the proven rate for generated properties decreased after the coverage agents' workflow, even though overall coverage increased." The coverage number went up while the fraction of properties that could actually be proven went down — and the published results ran with no human review in the loop. That is the metric improving while assurance does not, measured and self-reported.

The sharpest 2026 datapoint on where AI stimulus generation actually stops comes from a hole-by-hole taxonomy of everything an agentic flow failed to close across 19 designs (arXiv:2604.15657). Under 7% of the residual holes were genuinely unhittable — tied-off integration logic, defensive dead code, infeasible boundaries. About 92% were reasoning frontiers: multi-module pipeline warm-up sequences (49.9% of the frontier bucket) and protocol sequencing (40.2%) — holes that require building a Wishbone burst model or an MDIO responder to reach. The authors' key finding: "the agent correctly diagnoses these problems, [but] it fails to implement the solutions." The agent can read a coverage report and explain the hole like a staff engineer; it cannot yet write the protocol machinery to hit it. Diagnosis has been automated. Generation, at protocol depth, has not.

So the reformulated claim that survives all the evidence: AI earns its keep on tail bins that are reachable and correctly specified. The unreachable and the unenumerated stay yours.

Formal UNR: The Workhorse With Its Own Wall

The least glamorous automation in this story predates the AI wave and out-delivers most of it. Unreachability analysis takes your partial coverage database plus the RTL, formally proves which uncovered targets cannot be hit under any stimulus, and emits an exclusion file back into your coverage flow — Jasper's UNR app, VC Formal's FCA invoked natively from the VCS shell, Questa CoverCheck. A Questa team's DVCon tutorial reported a PCIe-bench UNR run of three hours that saved an estimated three weeks of manual analysis; Synopsys's blog cites Cisco seeing a 9% coverage improvement from pruning noise. One design in that same tutorial had over 3,000 unreachable coverage elements — at even 15 minutes of human triage each, that is 4.5 person-months of analysis a formal engine did before lunch.

Two things keep this section honest. First, nobody has a consistent answer to "how much does UNR save": Synopsys's own materials claim 40–80% verification-effort savings in one publication and 8–80% in another. The technique is real; the aggregate number is marketing. Second, a Qualcomm engineer's account from a VC Formal SIG punctures the assumption that UNR is a solved deployment: on their larger configuration — over 200 million coverage goals — the analysis produces claims at a scale where "millions… may need manual review by design experts, an impractical expectation only manageable through engineer-defined blanket exceptions which undermine the integrity of the analysis." And the exceptions don't port across configurations or even successive RTL drops. UNR "still is not as mainstream as you might imagine."

One reframe worth stealing from Synopsys's FCA documentation: an unreachable target you expected to be reachable is not a waiver candidate — it is a design bug wearing a coverage costume. And note that UNR-versus-ML is a false rivalry: NVIDIA's published flow runs test grading, UNR, and VSO.ai together. The formal engine prunes the impossible; the ML steers toward the merely improbable.

Exclusions Are Code

Every closure flow — manual, formal, or AI — terminates in the same artifact: an exclusion list that redefines what 100% means. Treat that artifact with the same rigor as RTL, because it carries the same tapeout risk. The best public, enforced example is OpenTitan's DV methodology:

  • Every exclusion carries a standardized annotation prefix — UNR, NON_RTL, UNSUPPORTED, EXTERNAL, LOW_RISK — making the exclusion base greppable and auditable. LOW_RISK is the honest bucket: a named, reviewed "we chose not to chase this," instead of a silent one.
  • Designers sign off on exclusions in PR review — the person who wrote the RTL certifies the code is genuinely uncoverable, not the person whose schedule benefits from the waiver.
  • "If any RTL changes happen to the design after the coverage exclusion file has been created, it needs to be redone and re-reviewed." Exclusion work starts only after design freeze, precisely because of this rule. The Questa tutorial gave the failure mode a name worth adopting: waiver rot — manually generated waivers "have to be maintained as the code changes."
  • Coverage is one line item among many: the V3 signoff gate also requires all assertions proven, no unreachable properties, and a nightly regression 100% passing with a week of soak. Closure is a portfolio, not a number.

The AI angle lands directly here. At least one new-entrant "coverage agent" product's headline capability is recommending exclusions with supporting evidence. Read that plainly: it closes coverage by removing targets. That may be exactly right — a good UNR-plus-triage assistant is valuable — but an AI-recommended exclusion must enter the same gate as a human one: annotated, designer-signed, invalidated on RTL change. An agent that can edit your exclusion file has write access to the definition of done.

Build Your Own Loop: The API Surfaces That Matter

Suppose you want the loop the papers describe — a model reading coverage, choosing what to run next — without waiting for a product. What can you actually build against, today? The answer has a clean structure, and it starts with zero APIs at all.

The loop you can write this afternoon lives entirely inside IEEE 1800. Every covergroup, coverpoint, and cross has a get_coverage() method, and pre_randomize() is the standard-blessed hook that runs before every randomize() call. Put them together and the testbench biases its own stimulus toward whatever is least covered:

function real calc_weight(opcode_t op);
  real cov;
  case (op)
    nop_op:  cov = covunit.cg.op_nop.get_coverage();
    load_op: cov = covunit.cg.op_load.get_coverage();
  endcase
  return (100 - cov) * 0.5;   // colder coverpoint => heavier weight
endfunction

function void pre_randomize();
  weight_nop  += calc_weight(nop_op);
  weight_load += calc_weight(load_op);
endfunction

No vendor dependency, works on all three simulators, and it is a real coverage-feedback controller — a proportional controller, in control-theory terms. It also teaches you the standard's load-bearing limitation by running into it: the LRM gives you coverage percentages, never bin identities. §19.9 defines exactly three covergroup system tasks ($set_coverage_db_name, $load_coverage_db, $get_coverage) plus the get_coverage methods; there is no standard way to ask which bins are empty from inside a running simulation. Everything more ambitious than proportional weighting needs the coverage database — which means it happens between runs, not during them.

The coverage database is the real interface, and one vendor documents it. Siemens publishes the full Questa UCDB C API — a 200+ page reference with worked examples shipped in the install — and the UCDB test record turns out to already be an ML training record: it stores the test name, the seed (verbatim from -sv_seed), a command field the docs describe as capturing "knob settings for parameterizable tests," CPU and simulation time, and pass/fail status. Stimulus knobs, seed, cost, outcome — the (action, cost, reward) tuple, persisted by the tool you already run, surviving merges. The Tcl layer above it is equally direct: vcover merge -testassociated (nothing per-test works without it), then coverage analyze -select cover -eq 0 — hole extraction as a query — then coverage ranktest, which emits ranktest.contrib and ranktest.noncontrib: machine-parseable lists of contributing and redundant tests, which is to say, labeled training data for a compression model, generated by a shipping tool. There is even a sanctioned place for an agent to stamp its own metadata into the database: coverage attribute -trendable.

The other two vendors gate their equivalents — Synopsys's coverage C API manual is marked Confidential; Cadence's IMC reference and vManager REST API live behind support portals. The practical trick: open-source consumers are the real documentation. OpenTitan's production fpv.tcl is a better JasperGold coverage reference than anything public from Cadence (check_cov -init before design load, -measure with a time limit, -report to parseable output — and waivers are just a Tcl file, which means an agent's exclusions are just a file it writes, subject to the previous section's gate). The Jenkins vManager plugin documents the vAPI REST surface by using it. And imc -execcmd "help report" makes the tool document itself.

On machine-readable output, the ground truth is humbling. The VCS manual documents exactly two urg output formats — text and HTML; zero occurrences of XML, JSON, or CSV. Cadence's flow, per OpenTitan's production scripts, emits text summaries and HTML (plus a native rank command). Across the industry, the practical egress contract for coverage data is parsing text reports — which retroactively explains the most striking pattern in the published work: neither industrial AI-closure paper used a vendor API. Nokia shelled out to ExecMan and parsed results; Infineon's agents shell out to Jasper and parse reports. The standard that was supposed to fix this — Accellera UCIS — was ratified in 2012 and never revised; its working group is inactive, no vendor publicly documents a conformant UCIS shared library, and the most credible open implementation (pyucis) had to write its own. The working loops route around the standard: Infineon's ISCAS 2025 flow skipped database extraction entirely and had the PyVSC coverage callback in a PyUVM monitor append stimulus values and bin hit/miss flags to a CSV during the run. Their stated reason is the thesis of the open-source path: "PyUVM testbenches offer a significant advantage in data collection compared to SystemVerilog-UVM testbenches." Python's edge is not nicer constraints. It is data egress — the testbench already lives in the language the ML lives in.

For tools with no API at all, two bridge patterns cover everything. Inside the simulation: DPI-C plus a TCP socket — both published RL closure loops (a DQN closing a compression encoder's CAM bins, and Infineon's Gym-based agents) independently built the same bridge, SystemVerilog importing a DPI function whose C side talks to a Python agent over a socket. One discipline point if you try it: the foreign call blocks the simulator, so make decisions in pre_randomize() or between transactions — a model consulted inside a sampling path perturbs timing and breaks seed-stability. Outside the simulation: a Tcl socket listener running inside any Tcl-shelled tool (vsim, ucli, Jasper, IMC) with a thin MCP server outside — a pattern already demonstrated by a third-party Xcelium debug server, and generalizable to every tool in your flow. No vendor cooperation required.

And the economics lesson that decides whether any of it pays: Nokia's loop reduced simulated tests by 43% and still failed to improve wall-clock time — worst case, regression time doubled — because every iteration tore down and restarted the simulation environment. The integration point that determines whether an AI loop is economical is not the coverage API. It is whether you can keep a warm simulator between decisions. Design for that first.

One boundary to scope your ambitions honestly: the commercial Tier B tools plug into the constraint solver and randomization kernel — a layer of the stack with no public API on any simulator. You cannot build VSO.ai or SimAI from outside. Everything upstream of the solver (test selection, knob choice, seed allocation) and downstream of it (hole extraction, triage, ranking, property generation, exclusion drafting) is buildable today with what this section named.

Piloting Without Getting Burned

If the post has a single operating principle, it is Mike Bartley's: separate productivity gains from assurance gains. A tool that halves your regression bill has improved productivity; whether your verification got better is a different question with different evidence, and the Infineon proven-rate result shows how easily a coverage metric can rise while assurance falls. His pilot guardrails compose into a checklist that fits on an index card: bound the pilot's scope; define success in engineering terms before deployment; version your datasets with traceability to DUT revision — "if the regression environment cannot reliably map failures to DUT revision, scenario, configuration, and known bug state, the model will learn noise"; route generated artifacts (tests, properties, exclusions) through the same review path as human-authored ones; and keep sign-off authority human.

Add the deployment lessons the case studies paid for: deploy at early milestones, not late (NVIDIA's late-project trial underperformed; the vendor now says so). Expect results proportional to your regression history — the derivative-design effect is real, and a team with years of archives will see numbers a fresh project will not. And run the trial on your design with a defined baseline, because the independent-evaluation landscape is thin enough to be its own finding: for the most heavily marketed tool in this space, every public number still routes through the vendor, and the one independent multi-project study found the ML tool roughly tied with plain seed ranking on its headline metric. That is not a reason to skip these tools. It is a reason to measure them the way you would measure anything else you were about to trust with tapeout: on your data, against your baseline, with the metric and the goal kept honestly apart.

The last mile of coverage closure has always been where verification stops being mechanical and starts being judgment — deciding what the model should contain, what the constraints should allow, what the design can never do, and what you are willing to sign. The verified wins in this post are real, and none of them moved that boundary. They cleared the brush on the road to it. Walk the last stretch yourself, and know exactly where it starts.


This post is part of the AI for DV series. Claims above trace to the primary sources linked inline; the research corpus and verification notes live with the series. Related: the Practitioner Playbook and the AI Reading List.

The AI-for-DV Research Reading List (2024-2026)

This is the citation home for the AI page. Every research claim in the Practitioner Playbook and its deep-dive posts traces back to an entry here. The list is grouped by theme, spans 2024–2026, and gets updated as the field moves.

Two kinds of entries: papers you should actually read (marked in the Start Here box), and papers you should know exist so you can find them when a problem lands on your desk. Where a named system has no standalone paper link, the entry points at the survey that covers it.

Start Here — Three Papers

  • The ASPDAC 2026 survey (Surveys, below) — the best current map of LLM-assisted verification: assertion generation, testbench automation, and RTL debug in one method tree.
  • CVDP (Code Generation, below) — the 783-problem hardware benchmark behind the 34% pass@1 ceiling. Read it before believing any codegen demo.
  • The Prompt Report (Prompt & Context, below) — the systematic survey the practical prompting patterns are drawn from.

Surveys & Field Maps

  • LLM-Assisted Circuit Verification: A Comprehensive Survey (ASPDAC 2026) — the field map: SVA generation (prompting / RAG / training-based branches), testbench and test automation, automated RTL debugging, and collaborative verification frameworks.
  • arxiv 2512.23189 — The Dawn of Agentic EDA — three-tier method taxonomy (prompt-based, fine-tuned, multi-agent), the "unit-test fallacy" critique of module-level benchmarks, and the call for an Open Agentic EDA Standard.

Prompt & Context Engineering

  • Anthropic, Effective Context Engineering for AI Agents (2025) — the canonical industry essay on the shift.
  • arxiv 2406.06608 — The Prompt Report, 58 techniques, PRISMA-grounded.
  • arxiv 2402.07927 — Systematic Survey of Prompt Engineering.
  • arxiv 2407.12994 — Prompt Engineering Methods for NLP Tasks survey.
  • arxiv 2509.21361 — Maximum Effective Context Window.
  • arxiv 2603.04814 — Beyond the Context Window, fact-based memory vs. long-context.
  • arxiv 2601.01954 — Reporting LLM Prompting in Automated SE.

Pair Programming & Debugging

Code Generation

  • arxiv 2604.24621 — Evaluation of LLM-Based SE Tools.
  • arxiv 2503.01245 — LLMs for Code Generation Comprehensive Survey.
  • arxiv 2505.02133 — Multi-Agent Collaboration and Runtime Debugging.
  • arxiv 2506.14074 — CVDP, 783-problem hardware benchmark.
  • arxiv 2506.07945 — ProtocolLLM, SV testbench benchmark.

Assertions & Formal Verification

  • arxiv 2410.23299 — FVEval (NVIDIA) — the first comprehensive benchmark for LLMs on formal verification: NL-to-SVA translation and direct assertion suggestion from RTL, in graded tiers.
  • AssertLLM (ASPDAC 2025) — multi-LLM pipeline generating SVAs directly from full specification documents; reported 89% of generated assertions syntactically and functionally correct on its evaluated design.
  • AutoSVA2, ChIRAAG, LASSO, AssertionForge, Hybrid-NL2SVA — the prompting / RAG / fine-tuned branches of the SVA-generation method tree; see the ASPDAC 2026 survey above for the full map.

Coverage & Benchmarks

  • Revisiting VerilogEval (ACM TODAES) — a year of LLM progress on the canonical RTL-generation benchmark, extended to spec-to-RTL tasks and failure classification.
  • arxiv 2311.00176 — ChipNeMo (NVIDIA) — domain-adapted LLMs for chip design; the reference point for the fine-tuned-model route.
  • LLM4DV (FCCM 2025) and VerilogReader (LAD 2024) — coverage-directed stimulus generation with feedback loops on uncovered bins; both covered in the ASPDAC 2026 survey above.

Technical Debt

  • arxiv 2601.06266 — SATD in LLM Software, three new debt types.
  • arxiv 2507.03536 — ACE: Validated LLM Refactorings.
  • arxiv 2501.09888 — Automated SATD Repayment.

Agentic Design Patterns

  • arxiv 2601.12560 — Agentic AI Architectures & Evaluation.
  • arxiv 2510.09244 — Fundamentals of Building Autonomous LLM Agents.
  • arxiv 2604.00835 — Agentic Tool Use.
  • arxiv 2510.25445 — Agentic AI Survey.
  • arxiv 2604.27643 — HAVEN, UVM testbench synthesis.
  • arxiv 2504.19959 — UVM².
  • arxiv 2605.04704 — UVMarvel, subsystem-level UVM testbench construction.

Security & IP Safety

  • arxiv 2604.01572 — VTS survey of AI-assisted hardware security verification — the five-stage pipeline from asset identification through countermeasure reasoning; covers SV-LLM and SoCureLLM (HOST 2025).
  • arxiv 2503.13116 — IP leakage from fine-tuning on in-house Verilog; reports up to 46.52% of generated code similar to the source IP.
  • arxiv 2405.07061 — LLMs and the Future of Chip Design — security risks and building trust in AI silicon flows.

Limits & Pitfalls

  • arxiv 2411.09916 — "Should I Give Up Now?" LLM Pitfalls in SE.

Conference Proceedings Worth Browsing

  • DVCon US 2025 proceedings — agentic verification, coverage closure with AI, and formal + GenAI flows from practitioners.
  • DVCon Europe 2025 program — three full sessions on AI in verification, including an industry-practice paper on RL-driven coverage closure.

The Observability Triad for UVM: Logs, Metrics, and Traces Your Testbench Already Emits

The Structured Logging post turned the testbench log from prose into a queryable JSONL stream — one signal, made sharp. But a log is only one of three kinds of evidence your testbench produces on every run. Software engineering names the full set the observability triad: logs, metrics, and traces. Your UVM environment already computes all three — it counts transactions, it tracks queue depths, it threads items from sequencer to scoreboard. It just never writes two of them down. This post maps the triad onto the testbench you already have, teaches the debug motion that uses all three signals together, and builds the missing piece: a metrics collector that turns numbers you print once into curves you can query.

The Problem: One-Signal Debugging

The regression test dies at the six-hour wall-clock limit. The last line of a 400 MB log reads UVM_ERROR ... [WATCHDOG] test timeout, and it tells you nothing except that the simulation was still alive and getting nowhere when the clock ran out. It is 2 AM, the nightly is red, and the debug session begins the only way it can: grep -n ERROR run.log to find where the wheels came off, then tail -100000 run.log | less, reading backward from the corpse. Every DV engineer knows this loop. You scroll up through tens of thousands of lines looking for the moment the run turned, reconstructing a six-hour story from the one page of it that happened to print near the end.

Here is what stings about that loop: the testbench was never short on evidence. It counted every transaction that crossed every interface. It watched the depth of every scoreboard queue, every sequencer arbitration queue, every analysis FIFO. It timed every request from launch to completion. Three different kinds of evidence — how much, how deep, how long — were computed continuously, on every clock, for the entire six hours. And exactly one of them was ever written down. That one was prose: UVM_INFO, UVM_ERROR, a human sentence at a time, the weakest and least structured signal the run produced. The numbers that would actually locate the failure were calculated and thrown away.

So the log sits in front of you at 2 AM, and the questions that matter are precisely the ones it cannot answer:

  • Was throughput already sagging in the minutes before the watchdog fired, or did the whole world stop at once? A gradual slope and a cliff are different bugs — one is backpressure, the other is a deadlock — and the log flattens both into the same silence.
  • When did a scoreboard queue start growing without draining, and which queue? A depth that climbs and never recovers is the fingerprint of a dropped completion, but a depth is a number over time, and the log kept no number and no time.
  • Did this run diverge from the fifty that passed before the failure signature ever appeared? For hours this run was statistically indistinguishable from the passing population — then it was not; the log gives you the moment of death but not the moment of departure.
  • Where did the six hours actually go — which phase, which interface, which wait? Wall-clock vanished into some blocking call that never returned, and a flat stream of timestamped sentences will not tell you which one held the whole test hostage.

Read those four again and notice what they have in common: not one is a question about messages. Every one asks about a quantity over time — a rate, a depth, a divergence, a duration — or about the path a single transaction took through the environment. Those are metric questions and trace questions, and you are trying to answer them by grepping a log because the log is the only artifact the run left behind. It is the wrong instrument, aimed at the wrong signal, and no amount of grep skill fixes an instrument that never recorded the measurement.

And every one of those numbers was computed during the run. Throughput was implicit in the transaction count the monitor already kept. Queue depth was a field the scoreboard read on every push and pop. Per-request latency was a subtraction the driver could have done in its sleep. The values existed, live, in the testbench's own memory — and then the objects were garbage-collected and the values went with them. Nothing was missing at runtime. Everything was missing at debug time, because nothing was persisted.

The fix, then, is not a sharper grep and not a bigger tail. It is not even a better logging macro. The fix is to stop discarding two-thirds of the evidence — to collect the other two signals and write them down beside the log, in a form you can query after the run is dead. Software engineering has a name for the discipline of keeping all three, and a well-worn theory of what each one is for. That is where we go next.

The Triad: Three Signals from the SRE Canon

The discipline is observability, rooted in Google's Site Reliability Engineering book and named as three signals — the "three pillars" — by the observability community that built on it. Map them straight onto a testbench and the abstraction stops being abstract.

Logs are discrete, timestamped events — high detail, high volume, one line per thing that happened. They answer what happened here: this transaction, on this interface, at this nanosecond, carried this payload. Your UVM_INFO stream is already this signal, and the Structured Logging post already sharpened it.

Metrics are numeric time series — a name, a value, a timestamp, repeated on a cadence. They are cheap to store and trivial to compare across runs, and they answer how is it trending: is throughput sagging, is that queue growing, is p99 latency drifting from last night's. A metric is the number the log threw away.

Traces are causally linked event chains that share an ID — the story of one transaction as it moves through the environment. They answer where did this request go and where did it stall: sequencer to driver to DUT to monitor to scoreboard, with a timestamp at every hop, so a six-hour disappearance resolves into the one edge that took six hours.

Metrics carry more structure than a bare number, and the Prometheus type system is the vocabulary for it — because the type dictates the query. A counter is monotonically increasing; you never read it raw, you query its rate — transactions issued per microsecond is the derivative of a counter that only ever climbs. A gauge is an instantaneous level that moves both ways; you plot it as-is, because scoreboard queue depth right now is the answer. A histogram buckets a distribution; you query its percentiles, because per-transaction latency is meaningless as a mean and everything as a p50 and p99. Counters get deltas, gauges get raw plots, histograms get percentiles — pick the wrong type and the right query becomes impossible.

Which numbers do you actually collect? Two methodologies answer that, and both return in the next section's inventory. Tom Wilkie's RED method covers every request stream: Rate (how many requests per second), Errors (how many failed), Duration (the latency distribution). Point RED at a UVM interface and it becomes transactions issued, mismatches flagged, and completion latency — a counter, a counter, and a histogram, one of each.

Brendan Gregg's USE method covers every resource instead: Utilization (how busy), Saturation (how much queued work is waiting), Errors. Point USE at a scoreboard or a sequencer arbiter and it becomes occupancy, queue depth, and dropped items — gauge, gauge, counter. RED watches the flow; USE watches the thing the flow runs through, and between them they name most of what a testbench should measure.

One honest caveat, from Honeycomb's Charity Majors: the "three pillars" framing oversells the split. Logs, metrics, and traces are storage formats, not three separate kinds of truth — a single stream of wide, richly attributed events can derive all three, and treating them as three siloed backends is how you end up unable to pivot from one to another mid-debug. That critique lands even harder in DV, and in our favor. Simulation is deterministic and closed-world, and the JSONL sidecar from the Structured Logging post is already that single wide-event stream — one file, every record tagged. We are not building three backends. This post teaches three lenses on the one stream you already have.

The Mapping: What Each Signal Already Is in UVM

The theory is portable, but the payoff is specific: point each of the three signals at a UVM testbench and it lands on something already running. You are not bolting observability onto your environment. You are naming machinery that is already there and writing down two-thirds of what it produces. Here is the whole mapping on one page — the SWE incarnation, where it already lives in your testbench, and the one thing still missing.

SignalSWE incarnationAlready in your testbenchThe missing piece
Logsstdout, Splunk, ELKuvm_report_server, `uvm_infoStructure — solved in the Structured Logging post
MetricsPrometheus, DatadogCoverage %, scoreboard depths, txn counters, sim-rateNobody samples them over time — built below
TracesOpenTelemetry, Jaegerseq_item lifecycle: sequencer → driver → monitor → scoreboardA propagated ID — the next post

Read the table top to bottom and the shape of the post appears: one row is done, one row is built below, one row is teased. Take them in order.

Logs — already sharp

Logs are the row with no work left. The Structured Logging post already took the uvm_report_server stream — the same `uvm_info/`uvm_error messages you have written a thousand times — and turned it from prose into a JSONL sidecar: one JSON object per event, written to a file beside the run. The contract is five fixed fields on every record — run to name the run, ts for the simulation timestamp in nanoseconds, sev for severity, id for the report ID, and comp for the full component path — plus a msg string and whatever domain fields the event carries (an address, a transaction id, a burst length). That is the entire log signal, already structured, already queryable, already shipping. Nothing in this post revisits how to produce it; the depth lives in that post and this section defers to it wholesale.

What this post adds is the other two signals, and it writes them to that same file. A metric is not a new sidecar or a second backend — it is another line in the JSONL stream, distinguished by a single kind field. A metric record carries kind set to metric; a log record carries no kind at all, so the absence of the field is the tag that says "this is a log." One file, two record types, one query surface. That decision — one wide-event stream, not three siloed stores — is exactly the Honeycomb critique from the last section, honored by construction. The log row is closed; the metrics row is where the code starts.

Metrics — the numbers you already print once

Metrics are the meat of the mapping, because this is the signal your testbench computes most eagerly and persists least. The two methodologies from the last section tell you precisely which numbers to keep, and they partition the testbench cleanly: RED treats every interface agent as a service carrying a request stream, and USE treats every testbench structure as a resource that stream flows through.

Point RED at an interface agent. Rate is the transaction count the monitor already increments on every completed item — a counter, queried as issued-per-microsecond. Errors is the scoreboard's mismatch count, the counter that matters most because a single nonzero increment is a failing run — the error signal, already tallied on every compare. Duration is per-transaction latency from launch to completion, a histogram, and it is the leading indicator: latency starts drifting in the buckets long before the throughput counter visibly sags, so the duration signal warns you a run is sick while the rate signal still looks healthy. Rate, errors, duration — a counter, a counter, a histogram, one per interface, all three already living in fields the agent updates every clock.

Point USE at a structure. Utilization is sequencer occupancy — the fraction of time the arbiter is handing out items rather than idle — a gauge. Saturation is queue depth: the scoreboard's pending-compare queue, the analysis FIFO, the response queue, read on every push and pop. That gauge is the one that predicts death — a depth that climbs and never drains is the fingerprint of a dropped completion, visible as a rising curve minutes before the watchdog fires. Errors is the uvm_error count, the same tally the report server keeps. Occupancy, depth, error count — gauge, gauge, counter, one set per resource.

Two more numbers belong to the run as a whole rather than to any single interface or structure. Functional coverage — $get_coverage(), or per-covergroup get_coverage() — is a gauge climbing toward 100, the metric that says whether the run is doing new work or spinning. And sim-rate, the ratio of Δsim-time to Δwall-clock, paired with the monitor's txn-rate counter, is what tells apart two failures a log renders identically. The clock generator does not know or care that the DUT is stuck, so a protocol-level hang leaves sim-rate normal — clocks toggling, sim-time advancing — while the txn-rate counter goes flat, nothing completing. A collapsed sim-rate is a different animal: the simulator itself is struggling — an X-storm, runaway event activity, host memory thrashing — and the DUT is incidental to it. Neither is visible from the message stream; both are one subtraction away from numbers the run already has.

Here is the observation the whole section turns on. Every metric named above — every rate, depth, occupancy, latency, coverage percentage — already exists in your testbench today, as an instantaneous value printed exactly once, in the final report, in the report_phase, as the run winds down. Your scoreboard already prints "12,481 transactions compared, 0 mismatches, max queue depth 34." A metric is that identical number, sampled on an interval instead of totalled at the end. The difference between "prints a total when it dies" and "emits a curve you can replay" is not new measurement — the measurement is done — it is a sampling loop and a writer, roughly 150 lines, built in Build Your Own: The Metrics Collector further down. The values are already in memory. The only missing act is writing them down on a cadence.

Traces — the chain that already exists

Traces are the row this post only teases, because the causal chain is already there in your environment, just implicit. A seq_item is born in the sequence, arbitrated by the sequencer, driven onto the bus by the driver, observed by the monitor, and compared in the scoreboard — sequencer → driver → monitor → scoreboard, a fixed lifecycle every transaction walks. That path is a trace in everything but name. What it lacks is a single propagated ID that survives every hop, so that all the records belonging to one transaction can be pulled out of the stream and laid end to end. The Structured Logging post already threads a txn_id field through its records for exactly this reason; a trace is what you get when that id becomes the spine of a query.

Give the chain a propagated ID and the six-hour-disappearance question from the top of the post becomes answerable directly: follow one transaction across all four components, measure the wall-clock or sim-time spent on each hop, and the one edge that swallowed the run resolves out of the noise. That is latency-per-hop, causally attributed, no grepping required. Building it — the ID scheme, the propagation through the sequence layering, the span-per-hop model — is more than a subsection can hold. Trace IDs and Context Propagation gets its own post; it is the next card in the Foundations row. For now the mapping is enough: the trace, like the metric, is a signal your testbench already emits and simply never writes down.

The Pivot: How the Three Work Together

The three signals are not three dashboards you check in parallel — they are one motion you run in sequence, a loop that walks a metric anomaly down to a root cause: a curve bends, that names a time window; the window plus a txn_id grep names the actor that stalled; the actor's log slice names what it said when it stalled; and if the root cause is not yet in hand, the new fact points you back at a different curve. Software engineers run this loop every time a pager goes off — dashboard to trace to log, narrowing at every hop — and the mapping in the last section is what makes it run on a testbench. Each signal answers a question the previous one could not, and hands the next one a filter.

flowchart LR
    M["📈 METRICS
which curve bent, and when?"] -->|"narrows to a time window"| T["🔗 TRACES
which transaction stalled, where?"] T -->|"identifies the actor"| L["📋 LOGS
what exactly did it say?"] L -->|"new hypothesis"| M

Watch it work on a real failure — a response-starvation hang, the kind that grep renders as six hours of silence and one dying gasp. Here is what the three signals recorded, laid on a single timeline:

Sim timeSignalEvidence
28 µshistogramp99 response latency starts creeping: 180 ns → 900 ns
32 µsgaugescoreboard expected-queue depth begins monotonic growth
36 µscountertxn completion rate flatlines at zero
40 µslog[WATCHDOG] test timeout — the only line grep would have found

Read the first column and the whole argument of this post is in the gap between the top row and the bottom one. The failure began at 28 µs. The symptom — the only line the log ever wrote — landed at 40 µs. Twelve microseconds separate the moment the run turned from the moment it announced it turned, and in a message-only workflow those twelve microseconds do not exist. Grep-from-the-end starts at 40 µs, at the watchdog line, and reads backward through prose that was written while everything already looked fine on the surface, hunting for a transition that left no distinctive sentence because the failure was numeric, not textual. That is the search that takes hours: you are reconstructing a slope from a stream that only ever recorded events, scrolling up through thousands of UVM_INFO lines that each say a normal thing happened.

The metrics reader starts somewhere else entirely. The p99 latency histogram — the leading indicator from the mapping, the signal that drifts in the buckets before the rate counter visibly sags — bent at 28 µs, and it bent on the response path specifically. So before opening a single log line, the reader has two things grep never gives you: a named suspect, the response stream, and a bounded window, roughly 28 to 40 µs. The saturation gauge confirms and localizes it — the scoreboard's expected-queue depth starts its monotonic climb at 32 µs, the fingerprint of completions that are launched and never returned, and it names the scoreboard partner as the structure filling up. Two queries in, the shape of the bug is already "the responder stopped feeding the scoreboard," and the window has tightened to six microseconds.

Now the pivot. The metrics have named the where and the when; they cannot say why, because a curve is a shape and not a sentence. So you cross from the metric stream to the log stream through the one field that spans both — the timestamp — and slice the JSONL sidecar to exactly the window the curves drew:

jq 'select(.kind != "metric" and .ts > 30000 and .ts < 36000)' run.sidecar.jsonl  # ts is in ns — this is the 30–36 µs window

That is not six hours of log. It is the ~20 lines that fall inside the bent window, and read in order they tell the story the watchdog line could not: the request stream keeps going — req after req issued, monitored, accepted, the outbound side perfectly healthy — while the response stream thins, one returned completion, then a longer gap, then another, then nothing. Requests in, responses trailing off to zero, the scoreboard queue swelling behind them. Twenty lines, read once, and the mechanism is plain: the responder model has stopped producing responses while the driver keeps issuing requests. One more look at the responder's own log lines in that slice and the cause is explicit — its available-credit count reached zero and never recovered. A credit leak in the responder model: it decremented on each request but missed the return path on some completion, bled its credits to nothing, and stalled. The DUT was fine. The testbench starved itself.

Count the work. Same bug, two workflows. Log-only: open a 400 MB file at its tail, grep for the error, then read backward for an hour or more reconstructing a numeric failure from a textual record that never named it, guessing at the window because nothing marks it. Triad: one histogram query names the suspect and the window, one gauge query names the partner and confirms the mechanism, one jq slice reads the twenty lines that explain it. Three queries against one file, a few minutes, and the pivot from which curve bent to what exactly did it say is a single shared ts. The evidence was always there; the difference is that two-thirds of it got written down, and that the writing-down turns a backward scroll into a forward filter.

This is Dave Agans' third rule of debugging, "Quit Thinking and Look," with instruments instead of vibes. Agans' warning is that engineers burn hours theorizing about what might be wrong instead of observing what is wrong — and the honest reason DV engineers theorize is that looking used to mean scrolling a log, which is slow and rarely conclusive, so guessing felt faster. The triad closes that gap. Looking is now three cheap queries that converge on a six-microsecond window, and only then do you open the waveform. The waveform is still the ground truth — that is where a DV debug ends — but a waveform is six hours wide and you cannot stare at all of it. The triad's entire job is to tell you which six microseconds to open. Metrics say when, traces say which transaction, logs say what it said, and the viewer, aimed at last, says why in silicon.

Build Your Own: The Metrics Collector

The mapping section promised the missing piece was small — a sampling loop and a writer, roughly 150 lines, standing between "prints a total when it dies" and "emits a curve you can replay." Here it is, and the line count holds. But the size is the least interesting thing about it. What matters is that the design obeys three rules borrowed straight from the software side of the house, and every one of them is load-bearing. The collector must be passive: it reads the testbench, it never steers it — read-only, zero-time, no objections raised, and critically no RNG touches, because a stray $urandom inside a probe perturbs the calling thread's random state and silently breaks seed-stability of the very run you are trying to observe — read-only with respect to the testbench; a probe may keep its own bookkeeping, a delta baseline or a histogram window, but never touches the object it observes. It must be pluggable: adding a fifth metric is one new class and zero edits to the collector, or the abstraction has failed. And it must be composable: every sample lands in the same JSONL sidecar as the log records from the Structured Logging post, one wide-event stream, with the kind field doing the discrimination. Passivity keeps the instrument from changing the experiment; pluggability keeps the instrument from rotting; composability keeps you on one query surface. Hold those three in mind — every design decision below is one of them, made concrete.

Start with the contract. The whole design inverts a dependency: the collector must not know what a throughput counter or a latency histogram is, only that it can be asked its name, its type, and its current value. SystemVerilog expresses that with an interface class — a pure interface, no implementation, no state — which is the language's answer to dependency inversion.

// The contract every metric source signs. `interface class` = a pure
// interface with no implementation — SV's answer to dependency inversion.
interface class metrics_probe;
  pure virtual function string name();         // e.g. "axi_wr_throughput"
  pure virtual function string mtype();        // "counter" | "gauge" | "histogram"
  pure virtual function string sample_json();  // value as a JSON fragment; MUST NOT touch testbench state
endclass

Three methods, and the third does the real work. sample_json() returns a JSON fragment, not a fixed scalar, and that is deliberate: a gauge can hand back "14" while a histogram hands back "{\"p50\":40,\"p95\":180,\"p99\":900,\"n\":512}", and the collector splices either one into the record without caring which it got. One contract, three metric types, no if (counter) ... else if (histogram) anywhere in the collector. If the phrase "every source signs a contract and the collector iterates over the signatories" sounds familiar, it should — this is the Observer pattern from the patterns series, where the subject holds a list of observers it knows only through their interface. The collector is the subject; the probes are the observers; the interface class is the abstraction that lets the two sides evolve independently.

Now the subject itself.

class tb_metrics_collector extends uvm_component;
  `uvm_component_utils(tb_metrics_collector)

  protected metrics_probe m_probes[$];
  protected time          m_interval = 1us;   // override: +metrics_interval_ns=<n>
  protected int           m_fd;
  protected string        m_run_id;

  function new(string name, uvm_component parent);
    super.new(name, parent);
  endfunction

  function void build_phase(uvm_phase phase);
    int unsigned iv_ns;
    if ($value$plusargs("metrics_interval_ns=%d", iv_ns)) m_interval = iv_ns * 1ns;
    if (!uvm_config_db#(string)::get(this, "", "run_id", m_run_id)) m_run_id = "0";
    m_fd = $fopen({get_full_name(), ".metrics.jsonl"}, "a");
  endfunction

  function void register_probe(metrics_probe p);
    m_probes.push_back(p);
  endfunction

  // Passive: no objection raised. The test ends when the test ends;
  // the collector just stops being scheduled.
  task run_phase(uvm_phase phase);
    forever begin
      #(m_interval);
      sample_all();
    end
  endtask

  protected function void sample_all();
    foreach (m_probes[i])
      $fdisplay(m_fd, $sformatf(
        "{\"run\":\"%s\",\"ts\":%0d,\"kind\":\"metric\",\"name\":\"%s\",\"mtype\":\"%s\",\"value\":%s}",
        m_run_id, $time, m_probes[i].name(), m_probes[i].mtype(),
        m_probes[i].sample_json()));
    $fflush(m_fd);
  endfunction

  // extract_phase is the phase DESIGNED for pulling data out of the
  // testbench — the last complete sample lands even on a clean finish.
  function void extract_phase(uvm_phase phase);
    sample_all();
    $fflush(m_fd);
  endfunction

  // Every sample is flushed as it lands (sample_all), so a $fatal or a grid
  // kill still leaves the last complete window on disk. final_phase just
  // closes the handle on a normal shutdown.
  function void final_phase(uvm_phase phase);
    $fclose(m_fd);
  endfunction
endclass

A few things earn their place. The interval is not hardcoded: build_phase reads +metrics_interval_ns=<n> off the command line and falls back to the config_db run_id so the record can be correlated across runs — sample cadence is a run-time knob, not a recompile. The file opens in append mode ("a"), which is what lets these metric lines coexist with log lines in one growing stream. In the format string, note %0d on $time rather than a stringified timestamp: it keeps the ts field a bare JSON number, so jq 'select(.ts > 30000)' — the exact query from the pivot section — compares numerically instead of choking on quotes. And look hard at run_phase: it raises no objection. That single absence is the passivity rule written in code. An objection would keep the test alive to finish sampling, which means the instrument would be extending the experiment — precisely the sin the passive rule forbids. Instead the collector rides extract_phase, the UVM phase designed for pulling data out of the testbench, so the last complete sample lands even on a clean finish, and because sample_all flushes every sample as it lands, even a $fatal or a wall-clock kill — deaths no UVM phase ever sees — still leaves the last complete window on disk; final_phase merely closes the handle on a normal shutdown. The test ends when the test ends; the collector just stops being scheduled. One last note for a single-sidecar setup: swap the $fopen for a handle shared with the report server from the Structured Logging post, and the metric lines and log lines pour into the identical stream — same file, kind discriminates.

That is the whole engine. Everything from here is a probe, and each probe is one small class that signs the contract. First, a counter — per-agent write throughput, delta-sampled.

// Counter: the monitor's analysis port increments; the collector's sample
// reads the delta since last sample. Rate = delta / interval.
class axi_throughput_probe extends uvm_subscriber #(axi_txn) implements metrics_probe;
  `uvm_component_utils(axi_throughput_probe)
  protected int unsigned m_count, m_last;

  function new(string name, uvm_component parent);
    super.new(name, parent);
  endfunction

  function void write(axi_txn t);
    m_count++;                        // hot path: one increment, nothing else
  endfunction

  virtual function string name();  return "axi_wr_throughput"; endfunction
  virtual function string mtype(); return "counter";           endfunction
  virtual function string sample_json();
    int unsigned delta = m_count - m_last;
    m_last = m_count;
    return $sformatf("%0d", delta);
  endfunction
endclass

The probe subscribes to the monitor's analysis port and its write does exactly one thing on the hot path — increment — because a subscriber that does real work per transaction taxes every run whether or not anyone reads the metric. The type discipline from the theory section shows up here: a counter is never emitted raw, so sample_json() returns the delta since the last sample and stashes the new baseline, handing the querier a per-interval count it can divide into a rate. One subtlety worth calling out for anyone reaching for their linter: metrics_probe::name() does not collide with uvm_object::get_name(). The contract method was named name() precisely so it sits alongside the UVM identity method rather than overriding it — the probe answers to both, and they mean different things.

Now a gauge, and here the implements keyword shows its power. The scoreboard does not need a separate probe object watching it from outside; the scoreboard is the probe. It already knows its own queue depth, so it signs the contract directly.

// Gauge: the scoreboard IS the probe. SV's `implements` keyword lets a
// component sign the metrics contract without inheriting anything new.
class axi_scoreboard extends uvm_scoreboard implements metrics_probe;
  `uvm_component_utils(axi_scoreboard)
  protected axi_txn m_expected[$];
  // ... normal scoreboard duties elided ...

  virtual function string name();  return "sb_depth"; endfunction
  virtual function string mtype(); return "gauge";    endfunction
  virtual function string sample_json();
    return $sformatf("%0d", m_expected.size());  // read-only: .size(), no pop
  endfunction
endclass

A gauge is an instantaneous level, so sample_json() just reports m_expected.size() as-is — no delta, no reset. And note what it does not do: it calls .size(), never .pop_front(). Reading the depth must not consume the queue, or the instrument would be altering the thing it measures. This is the same saturation gauge that climbed at 32 µs in the pivot timeline — the fingerprint of completions launched and never returned — now emitted on a cadence instead of printed once as "max queue depth 34."

The third type is the histogram, and it is the one the pivot section leaned on hardest.

// Histogram: monitors report request->response latency; the probe sorts
// the current window at sample time and emits percentiles, then resets.
class axi_latency_probe extends uvm_subscriber #(axi_txn) implements metrics_probe;
  `uvm_component_utils(axi_latency_probe)
  protected time m_window[$];

  function new(string name, uvm_component parent);
    super.new(name, parent);
  endfunction

  function void write(axi_txn t);
    m_window.push_back(t.rsp_time - t.req_time);
  endfunction

  virtual function string name();  return "axi_rd_latency"; endfunction
  virtual function string mtype(); return "histogram";      endfunction
  virtual function string sample_json();
    time sorted[$] = m_window; string s;
    if (sorted.size() == 0) return "{\"n\":0}";
    sorted.sort();
    s = $sformatf("{\"p50\":%0d,\"p95\":%0d,\"p99\":%0d,\"n\":%0d}",
                  sorted[(sorted.size()-1)*50/100],   // (n-1)*p/100: always in range
                  sorted[(sorted.size()-1)*95/100],
                  sorted[(sorted.size()-1)*99/100],
                  sorted.size());
    m_window.delete();
    return s;
  endfunction
endclass

The write path stays cheap — one subtraction pushed onto a window — and the expensive step, the sort, happens only at sample time when someone is actually going to read the number. The percentile index (n-1)*p/100 is chosen so it always lands inside the array for any non-empty window, which is why the size-zero case returns early with {"n":0}. Then the window is cleared, so each sample describes the latencies in that interval rather than the run to date. This is the leading indicator made concrete: p99 on the response path is the signal that bent at 28 µs in the pivot timeline — a full 8 µs before the throughput counter flatlined at 36. The histogram warns you a run is sick while the counter still looks healthy, and now that warning is a queryable curve.

Three metric types, three probes, and every one of them lived entirely inside the HDL. The fourth reaches for something SystemVerilog does not have. Sim-rate is Δsim-time over Δwall-clock, and the language has no wall-clock — $time is simulated nanoseconds, and $system() shells out for a status code, not a value you can read back. This is the moment to open the software toolbox and pull out DPI-C. Three lines of C give the testbench a monotonic wall clock.

// SystemVerilog has no wall-clock. $system() returns a status, not a value.
// Three lines of C fix that:
import "DPI-C" function longint tb_epoch_ms();
// tb_epoch.c — compile into the sim with your tool's DPI flow
#include <time.h>
long long tb_epoch_ms(void) {
  struct timespec ts;
  clock_gettime(CLOCK_MONOTONIC, &ts);
  return (long long)ts.tv_sec * 1000 + ts.tv_nsec / 1000000;
}

With a wall clock imported, the probe is another gauge.

// Gauge: sim-nanoseconds advanced per wall-millisecond. Read WITH the
// txn-rate counter: txn rate flat while sim_rate stays normal = protocol
// hang (clocks still toggling, nothing completing); sim_rate collapsed =
// the simulator itself is struggling (X-storm, runaway event activity,
// host memory thrash).
class sim_rate_probe implements metrics_probe;
  protected time    m_last_sim;
  protected longint m_last_wall;

  virtual function string name();  return "sim_rate"; endfunction
  virtual function string mtype(); return "gauge";    endfunction
  virtual function string sample_json();
    time    now_sim  = $time;
    longint now_wall = tb_epoch_ms();
    real    rate     = (now_wall == m_last_wall) ? 0.0
                     : real'(now_sim - m_last_sim) / real'(now_wall - m_last_wall);
    m_last_sim  = now_sim;
    m_last_wall = now_wall;
    return $sformatf("%.1f", rate);
  endfunction
endclass

The comment carries the diagnostic payload, and it is the one from the mapping section, unchanged: read sim-rate with the txn-rate counter and the pair separates two failures a log renders identically. Txn rate flat while sim-rate stays normal is a protocol hang — clocks toggling, sim-time advancing, nothing completing. Sim-rate collapsed is a different animal — the simulator itself is struggling under an X-storm, runaway event activity, or host memory thrash, and the DUT is incidental. One more structural note: sim_rate_probe extends nothing. It is a plain class, not a uvm_component, because a probe needs only to sign the contract — the collector holds it by its metrics_probe handle and never asks it to be anything more. That is the dependency inversion paying off; the collector's list does not care what class of object it holds. And with four probes in hand, the fifth writes itself: a coverage_pct gauge — $get_coverage() wrapped in the same three-method shape — is the reader's first exercise; it is six lines.

Which leaves the wiring, and the wiring is where the whole design either earns its keep or does not.

function void tb_env::connect_phase(uvm_phase phase);
  super.connect_phase(phase);
  axi_agent.mon.ap.connect(m_wr_tput.analysis_export);
  axi_agent.mon.ap.connect(m_rd_lat.analysis_export);
  m_metrics.register_probe(m_wr_tput);   // counter
  m_metrics.register_probe(m_rd_lat);    // histogram
  m_metrics.register_probe(m_sb);        // gauge — the scoreboard itself
  m_metrics.register_probe(m_sim_rate);  // gauge — plain class
endfunction

The subscribers hook onto the monitor's analysis port, and then all four probes register with the collector through the same one-line register_probe call — a counter, a histogram, and two gauges, one of which is the scoreboard measuring itself and one of which is a plain non-component class, handled identically because the collector sees only the contract. This is the pluggability rule collecting its debt: adding metric #5, that coverage_pct gauge, means writing one class and adding one register_probe line here — it touches zero lines inside tb_metrics_collector, because the collector was never told what any of its probes are. That is dependency inversion earning its keep, and it is why the whole collector fits in 150 lines while staying open to every metric your testbench will ever want to emit.

Querying the Signals

Everything in the pivot section ran on one query primitive: jq, pointed at the JSONL sidecar, filtering on kind and ts. The filenames below are illustrative, not prescriptive — a metrics sidecar and a log sidecar, shown here as two files. Section 5's collector writes wherever you point its $fopen: one shared stream discriminated by kind, or two files that cat concatenates into one — the queries below run unchanged either way. That is not a coincidence to gloss over — it is the entire payoff of writing one wide-event stream instead of three siloed backends. You do not need a metrics database, a trace viewer, or a query language with a learning curve. You need four one-liners, and you already have three of them memorized from the sections above. Here they are as a toolkit, each doing one job, none of them needing anything more exotic than the jq already on your workstation.

The first turns a single metric into a plottable series — pick a name, drop everything else, keep timestamp and value:

# One metric as CSV — feed to any plotter
jq -r 'select(.kind=="metric" and .name=="sb_depth") | [.ts, .value] | @csv' run.metrics.jsonl

Against the fixture built for this section — sb_depth sampled every microsecond, climbing from 32 µs — this line reads out 0,1, 1000,0, 2000,1, in order, ready for any plotter that takes CSV.

The second answers a narrower question — not the whole series, but the one instant that matters: where does the counter last read nonzero, the sample just before the flatline the pivot section built its whole case around.

# Hang window: last sample where throughput was still nonzero
jq -r 'select(.name=="axi_wr_throughput" and (.value|tonumber) > 0) | .ts' run.metrics.jsonl | tail -1

On the same fixture — throughput steady at 8 through 35 µs, zero from 36 µs on — this returns 35000: one number, ns, and it is the right edge of the window you need to slice next.

The third is the one line that reaches straight into a histogram's structure rather than a scalar, because the percentile is stored inline on every sample and needs no recomputation at query time:

# p99 latency curve — histograms carry their percentiles inline
jq -r 'select(.name=="axi_rd_latency") | [.ts, .value.p99] | @csv' run.metrics.jsonl

Run against the fixture, the p99 column sits flat at 180 through 28 µs and then climbs — 230, 280, 330 — exactly the drift the pivot's histogram query depended on, read straight off .value.p99 with no post-processing.

The fourth is the pivot itself, restated as a reusable line rather than a one-off: metrics named the window, now slice the log stream — the records with no kind field — to that same span.

# The pivot: metrics found the window, now slice the LOG stream to it
jq 'select(.kind != "metric" and .ts > 30000 and .ts < 36000)' run.sidecar.jsonl

Four lines, one file format, no new tool. Piped to a CSV, a metric series is one step from a picture, and a picture is what makes a bent curve visible at a glance instead of buried in a column of numbers. Feed two series through gnuplot and the throughput drop and the depth climb from the pivot's timeline land on one plot, one axis each:

set datafile separator ","
set y2tics
plot "tput.csv"  using 1:2 with lines title "wr throughput",      \
     "depth.csv" using 1:2 with lines title "sb depth" axes x1y2

One PNG per failing run, generated in the regression triage script — the curve is the first debug step.

The real payoff is not one run's plot — it is fifty of them. Run the same jq → CSV pipeline over every passing run from the last week, overlay all fifty curves on the failing run's curve on one axis, same metric, same color for the passing band, one line standing out in the failing run's color. Fifty healthy throughput traces sit in a tight band; the failing run's trace rides inside that band for hours and then peels away from it, and the instant it peels away is not a guess — it is a pixel you can point at. That instant is the debug starting line. It is earlier than the watchdog line by however many microseconds the pivot section found, and it is exactly the moment worth asking "what changed here" about, because the answer to that question is the bug.

Fifty passing runs and one failing run, overlaid, is one query away — but ten thousand runs a night, with dozens of distinct failure signatures tangled together, is a different problem than eyeballing one divergent line on one plot. That is where this series goes next: the Triage category picks up from here, with an Error Fingerprinting card asking what pattern a failure leaves behind when you cannot look at each of ten thousand plots by hand. The signals are captured; querying them at that scale is the next post's problem, not this one's.

Quick Reference

Everything above compresses to one table, one checklist, and a reading list. Bookmark this section; it is the part you come back to at 2 AM instead of re-reading the whole post.

SignalUVM sourceCollection costFirst question it answers
Logsuvm_report_server → JSONL sidecarDone — previous postWhat exactly happened at time T?
Metrics (counter)analysis-port subscriber, delta-sampled~20 lines/probeWas it still making progress?
Metrics (gauge)scoreboard/queue .size() via implements~10 lines/probeWhat was filling up, and when did it start?
Metrics (histogram)latency window, percentiles at sample time~30 lines/probeWas it getting slower before it stopped?
Tracestxn_id threading (next post)One UUID + disciplineWhere did THIS transaction stall?

Adopt the triad in the order that pays off fastest, not the order this post presented it:

  1. Drop in tb_metrics_collector with the sim_rate probe — zero testbench coupling, and your regression can already tell a hung DUT from a thrashing simulator.
  2. Register scoreboard-depth (implements metrics_probe, ~10 lines) and one per-agent throughput probe.
  3. Add the jq → gnuplot step to the fail-triage script: every failing run gets its curves rendered before a human looks at it.

Each step stands alone — step 1 alone already answers the "is the simulator or the DUT stuck" question from the mapping section, without touching a single scoreboard or agent class.

Further reading:

  • Google, Site Reliability Engineering — the chapter on monitoring distributed systems, the monitoring foundations under the three-signal framing this post maps onto UVM.
  • Peter Bourgon — Metrics, Tracing, and Logging — the post that named the three pillars.
  • Charity Majors (Honeycomb) — the "three pillars" critique: logs, metrics, and traces are storage formats, not separate truths.
  • Brendan Gregg — the USE method (Utilization, Saturation, Errors), for naming what to measure on any resource.
  • Tom Wilkie — the RED method (Rate, Errors, Duration), for naming what to measure on any request stream.
  • Prometheus documentation — metric types (counter, gauge, histogram), the vocabulary this post borrows wholesale.
  • OpenTelemetry — semantic conventions, the closest thing the industry has to a standard schema for traces.
  • David Agans, Debugging — Rule 3, "Quit Thinking and Look," the rule the pivot section runs on.

That closes the loop this post opened. The Structured Logging post gave the testbench its first sharp signal; this one added the two it was throwing away. Next in Foundations: Trace IDs & Context Propagation — the UUID that turns your log stream into traces, and the piece this post kept teasing and deferring. Until then, everything here — logs, metrics, and the queries that connect them — lives on the Debug hub, alongside the rest of the triage toolkit this series is building one card at a time.