Error Fingerprinting for UVM Regressions: Bucketing 10,000 Failures the Way Sentry and Windows Do

The overnight regression finished with 10,234 failures. The structured-logging post ended with a one-line jq that collapses them into a dozen buckets by hashing three fields. That line works, and it is also where most teams stop. Then, two weeks later, the buckets start lying: a scoreboard message that hides five different bugs, an address in a message that splits one bug into four hundred buckets, a hang that lands in whichever bucket the watchdog message happens to match. Nobody trusts the dashboard, and the senior engineer is back to opening logs in tabs.

The software world hit exactly this wall twenty-five years ago and wrote down what it learned. Microsoft's Windows Error Reporting has bucketed crash reports from a billion machines since 1999. Sentry and Rollbar bucket exceptions for hundreds of thousands of applications and expose their grouping rules in public documentation. Microsoft Research published the algorithm that replaced message matching with call-stack matching. This post takes those results and rebuilds error fingerprinting for UVM regressions on top of them: what a fingerprint must be orthogonal to, what the "stack trace" of a testbench failure actually is, how to normalise a message without destroying its signal, and how to keep the rules honest as the project changes.

Two ways a fingerprint lies

The Windows Error Reporting team defined the goal in one sentence: a bucketing algorithm should maintain orthogonality, one bug per bucket and one bucket per bug. Every failure of a fingerprint is a failure of one half of that sentence.

  • Under-split: two bugs in one bucket. Your scoreboard reports SCB_MISMATCH for every data error it sees. A parity bug in the write path and a byte-lane bug in the read path both surface as SCB_MISMATCH from the same component. Bucketed on message id and component, they are one bucket with 1,450 hits, one owner, and one very confused afternoon.
  • Over-split: one bug in many buckets. The same scoreboard prints the address in the message. One bug, hit by 400 seeds at 400 addresses, becomes 400 buckets of one hit each. The dashboard says "400 distinct problems" and the team stops reading it.

The Vennsa and University of Toronto paper that called failure triage "the neglected debugging problem" drew exactly these two pictures for hardware: two distinct bugs caught by the same checker, and one bug caught by different checkers because different stimulus propagated it along different paths. Their observation was that the industry's automation, where it existed, binned "purely on the error message and the owners of the failing tests", and that this was the source of both failure modes.

WER's vocabulary for fixing this is worth adopting as-is. A heuristic that adds information to the fingerprint is expanding: it increases the bucket count so that two bugs stop sharing one. A heuristic that removes information is condensing: it decreases the bucket count so that one bug stops spanning several. Their table of client-side heuristics is mostly expanding (add the module name, the offset, the exception code, the hang wait-chain root) with a few deliberate condensing ones (replace the module and offset with an in-code assert ID, because the assert ID identifies the bug better than where it fired). Every change you make to a fingerprint rule is one of these two moves, and you should know which one you are making and why.

DV translation The equivalent of WER's assert_tags condensing heuristic is using the SVA property name or the uvm_report id instead of the message text. The equivalent of the expanding hang_wait_chain heuristic is fingerprinting a timeout by which sequence was blocked and on what, not by the watchdog's generic message.

The failure stack: what your testbench's stack trace is

Sentry's default grouping does not look at the message first. It looks at the stack trace, and it distinguishes in-app frames, your code, from library and framework frames, which it ignores for grouping. Rollbar does the same: its exception fingerprint is a hash of the exception class plus the file and method names of the frames, with line numbers dropped and framework boilerplate frames removed. The 2012 ReBucket paper from Microsoft Research went further and replaced exact matching with a similarity measure: frames near the top of the stack (the crash point) weigh more than frames near the bottom, and two stacks that match the same functions at slightly different depths still count as similar. ReBucket also strips "immune functions", library code that is trusted enough to be unlikely to hold the bug.

A UVM failure has a stack trace too. It is just not in the place a software engineer would look. The frames, from top to bottom, are:

FrameSourceWeightWhy
Assertion or checker idrm.get_id(), property nameHighestClosest to the observation of the bug, the "crash point"
Reporting componentrm.get_report_object().get_full_name(), instance indices maskedHighWhich monitor or scoreboard saw it
Phaseuvm_phase at the time of the errorMediumReset-phase failures and run-phase failures are rarely the same bug
Running sequence chainget_parent_sequence() walked to the rootMedium, decaying downwardWhat stimulus was active; the outer virtual sequence matters less than the innermost
Test name+UVM_TESTNAMELowestThe frame most scripts bucket on, and the least informative: one test hits many bugs and one bug hits many tests

Notice the weight column is the inverse of common practice. Grouping by test name is grouping by the bottom frame, the main() of the failure. ReBucket's position-dependent model says that is the frame to trust least.

The in-app rule also transfers cleanly. Components in your environment are in-app. Anything under uvm_pkg and inside a purchased VIP is a library frame: keep it for context, drop it from the fingerprint. A failure reported by uvm_test_top.env.pcie_vip.dl_layer.crc_checker is fingerprinted at the VIP boundary, env.pcie_vip, plus the VIP's own error id, because the frames inside it are not yours to fix and their internal paths change with every VIP release.

Here is the collector. It runs inside a report catcher, so it sees every error at the moment it is raised, and it emits one compact record per failing simulation, on the first error only.

class failure_stack_catcher extends uvm_report_catcher;
  `uvm_object_utils(failure_stack_catcher)

  bit             captured;
  string          library_roots[$] = '{"pcie_vip", "axi_vip"};
  // Set from phase_started() in your base test: current_phase = phase.get_name();
  static string   current_phase = "build";

  function new(string name = "failure_stack_catcher");
    super.new(name);
  endfunction

  // Mask instance indices: env.axi_agent[3].monitor -> env.axi_agent[*].monitor
  function string mask_indices(string path);
    string out = "";
    foreach (path[i]) begin
      if (path[i] == "[") begin
        out = {out, "[*]"};
        while (i < path.len() && path[i] != "]") i++;
      end else out = {out, path[i]};
    end
    return out;
  endfunction

  // First index of sub in s, or -1 (SystemVerilog has no built-in substring search)
  function int index_of(string s, string sub);
    for (int i = 0; i + sub.len() <= s.len(); i++)
      if (s.substr(i, i + sub.len() - 1) == sub) return i;
    return -1;
  endfunction

  // Cut the path at a library boundary so VIP internals do not enter the fingerprint
  function string to_in_app(string path);
    foreach (library_roots[k]) begin
      int idx = index_of(path, library_roots[k]);
      if (idx >= 0) return path.substr(0, idx + library_roots[k].len() - 1);
    end
    return path;
  endfunction

  function action_e catch();
    uvm_report_object ro;
    uvm_sequence_base seq;
    string frames[$];
    if (captured || get_severity() < UVM_ERROR) return THROW;
    captured = 1;

    ro = get_client();
    frames.push_back({"id:", get_id()});
    frames.push_back({"comp:", mask_indices(to_in_app(ro.get_full_name()))});
    frames.push_back({"phase:", current_phase});

    // Sequence chain, innermost first, stored as kinds not instances
    if ($cast(seq, ro) == 0) seq = running_sequence(ro);
    while (seq != null) begin
      frames.push_back({"seq:", seq.get_type_name()});
      seq = seq.get_parent_sequence();
    end

    log_failure_record(frames, get_message());
    return THROW;
  endfunction
endclass

Three hooks are yours. Override phase_started() in the base test with one line, failure_stack_catcher::current_phase = phase.get_name();, so the catcher knows the phase without reaching into UVM internals. running_sequence() and log_failure_record() are the other two: the first resolves the sequence currently on the sequencer that drives the reporting agent, the second appends one JSON line to the sidecar file described in the structured-logging post. The record carries the frames, the raw message, and the run header fields (seed, test, RTL hash, tool version). Everything below consumes that record.

Normalising the message without losing the signal

Rollbar publishes its message-normalisation list, and it is short: strip dates, timestamps, email addresses, IP addresses, decimal numbers, integers of two or more digits, and hex values of four or more digits, then hash what is left. The interesting part is the exception: error codes and HTTP status codes are deliberately kept, because a 404 and a 500 in the same handler are different bugs. "Strip things that look like data, keep things that look like codes" is the whole rule.

The DV version of the list:

MaskPatternKeep or stripReason
Simulation time@ 12345 ns, time=...StripChanges with every seed
Seedseed=..., +ntb_random_seedStripRun identity, not bug identity
Addresses and data0x[0-9a-f]{4,}, 'h...StripThe classic over-split source
Transaction and packet idsid=NN, seq_id, tag=StripPer-run counters
Instance indices[3] in pathsStripSame bug across ports
Small integers0 to 9 standing aloneKeepUsually a lane, a channel, a state: a code, not data
State and opcode namesIDLE, L0s, WRITE_BURSTKeepThe DV equivalent of an error code
Checker verdict wordsexpected, actual, timeout, underflowKeepDistinguish mismatch from protocol violation

The one judgment call is the two-digit rule. Rollbar strips integers of two or more digits and keeps single digits. That heuristic is right for hardware too: a lane index of 3 is a code, a burst length of 256 is data. Apply the same rule and you get 90 percent of the benefit with no per-message configuration.

import re

MASKS = [
    (r"@\s*\d+\s*[pnuf]?s\b",              "@<T>"),    # sim time
    (r"\b(seed|ntb_random_seed)\s*[=:]\s*\d+", r"\1=<SEED>"),
    (r"0x[0-9a-fA-F]{4,}",                 "<HEX>"),    # addresses, data
    (r"'h[0-9a-fA-F_]{4,}",                "<HEX>"),
    (r"\b(id|tag|seq_id|txn)\s*[=:]\s*\d+", r"\1=<ID>"),
    (r"\[\d+\]",                           "[*]"),      # instance indices
    (r"\b\d{2,}\b",                        "<N>"),      # integers of 2+ digits
]

def normalise(msg: str) -> str:
    for pat, rep in MASKS:
        msg = re.sub(pat, rep, msg)
    return re.sub(r"\s+", " ", msg).strip()

# "SCB_MISMATCH addr=0x4000_1000 lane 3 expected 0xdead actual 0xbeef @ 128340 ns"
#   -> "SCB_MISMATCH addr=<HEX> lane 3 expected <HEX> actual <HEX> @<T>"

If your logs are prose because the testbench predates structured logging, you do not have to write masks by hand. The Drain algorithm, published in 2017 and maintained as the drain3 package, mines message templates from a stream of log lines: variable positions become <*> automatically, a similarity threshold decides when a line starts a new template, and custom masks for hex and integers are applied first. Feed it your UVM_ERROR lines and the template id it returns is a usable normalised message on day one.

Label in the simulation, classify in the script

WER's most transferable design decision is that bucketing happens in two phases in two places. Labeling runs on the client, from the evidence available at the moment of the crash, and is cheap. Classifying runs on the server, once more data has arrived (symbols, memory dumps, a second report from the same bucket), and can move a report to a better bucket. Bucketing is progressive: the first answer is not the final answer.

Your simulation is the client. Your triage script is the server.

  • Level 1, in the simulation: a label from immediate evidence only. Top two frames plus the normalised message: id|component|template. It costs nothing, needs no other run, and is right often enough to route the failure.
  • Level 2, in the triage script: a classification that adds evidence the simulation could not see cheaply or at all. The last five transaction kinds from the ring buffer. The DUT state signature. Whether the failure was preceded by a reset or a power transition. Whether the same seed passed on yesterday's RTL. This is where the scoreboard's five hidden bugs get separated, because the read-path bug always has a READ_BURST in the last five kinds and the write-path bug never does.

Two records that share a Level 1 label but split at Level 2 tell you the label was under-split, and that is an expanding heuristic waiting to be promoted into the simulation-side rule. Two Level 2 buckets that a human keeps merging tell you the opposite. The levels are not just a pipeline. They are how the rule set learns.

From the literature WER's !analyze tool grew to roughly 100,000 lines implementing about 500 bucketing heuristics, at about one new heuristic a week for a decade. That number is the honest forecast for any fingerprint rule set: it will keep changing. Store the rule-set version in every record, and never compare buckets across versions without re-bucketing the old records.

Rules, overrides and merges

Sentry and Rollbar both expose the same three controls, and their syntaxes are worth copying because they were arrived at by years of users asking for the same things.

Fingerprint rules match an event and replace the fingerprint. Sentry writes them as one line each, first match wins:

error.type:DatabaseUnavailable -> system-down
error.value:"connection error: *" -> connection-error, {{ transaction }}
logger:my.package.* level:error -> error-logger, {{ logger }} title="Error from Logger {{ logger }}"

Rollbar writes them as JSON with a condition, a fingerprint template and a title, and its condition language has the operators you would expect: equality, membership, prefix and suffix, numeric comparison, regex, combined with any, all and none. Both let a rule refer back to the default fingerprint, so a rule can refine grouping instead of replacing it.

Merges are the human override for under-split rules: two issues merged so that future events land in one, with an unmerge that restores the originals. Sentry now also proposes merges from a transformer embedding of stack traces, which is where "similar enough" grouping ends up once exact matching runs out of road.

For a regression flow the same three controls fit in one YAML file next to the testbench, evaluated by the triage script in order:

version: 7
rules:
  # Condensing: every VIP-internal CRC error is one bug at the VIP boundary
  - match: { id: "PCIE_DL_CRC*", comp: "env.pcie_vip*" }
    fingerprint: "pcie-vip-crc|{{ phase }}"
    title: "PCIe VIP CRC error"

  # Expanding: the scoreboard hides read-path and write-path bugs behind one id
  - match: { id: "SCB_MISMATCH", last_kinds: "*READ_BURST*" }
    fingerprint: "{{ default }}|read-path"

  # Timeouts are grouped by what was blocked, never by the watchdog text
  - match: { id: "WATCHDOG_TIMEOUT" }
    fingerprint: "hang|{{ seq[0] }}|{{ comp }}"
    title: "Hang in {{ seq[0] }}"

merges:
  - into: "a1f9c2e0"
    from: ["7b3d0e11", "c04e9a77"]
    reason: "Same byte-enable bug seen by monitor and scoreboard, JIRA DV-418"

The evaluator is short. Conditions are shell-style globs against record fields, templates pull fields into the fingerprint string, and the hash is taken over the final string so that rules can be reordered without changing existing bucket ids.

import fnmatch, hashlib, json, re, yaml

def matches(rule, rec):
    return all(fnmatch.fnmatch(str(rec.get(k, "")), pat)
               for k, pat in rule["match"].items())

def render(template, rec, default):
    def sub(m):
        key = m.group(1).strip()
        if key == "default":
            return default
        if key.startswith("seq["):
            return rec["seq"][int(key[4:-1])] if rec.get("seq") else ""
        return str(rec.get(key, ""))
    return re.sub(r"{{\s*(.*?)\s*}}", sub, template)

def fingerprint(rec, cfg):
    default = "|".join([rec["id"], rec["comp"], rec["template"]])   # Level 1 label
    fp, title = default, rec["id"]
    for rule in cfg["rules"]:
        if matches(rule, rec):
            fp = render(rule["fingerprint"], rec, default)
            title = render(rule.get("title", title), rec, default)
            break
    digest = hashlib.sha1(fp.encode()).hexdigest()[:8]
    for m in cfg.get("merges", []):
        if digest in m["from"]:
            digest = m["into"]
    return digest, fp, title, cfg["version"]

Three properties of this design matter more than the code. Rules live in version control next to the testbench, so a rule change is reviewed like any other change. The rule-set version travels with every record. And the merge table is data, not a special case, so an override made at 2 AM survives the next rule edit.

The first error is not always the root

Every fingerprint pipeline starts by taking the first error in the log. Two situations break that assumption, and both have a software precedent.

Hangs. WER's hang heuristic does not fingerprint a hang by the fact that the UI stopped. It walks the chain of threads waiting on synchronisation objects, starting from the input thread, to find the root of the wait chain, and buckets on that. The DV equivalent: a UVM_FATAL from the watchdog is the observation, not the bug. The fingerprint should come from the watchdog's structured report of what was blocked, the sequence that never got its response, the sequencer it was waiting on, and the phase, which is why the rule above fingerprints WATCHDOG_TIMEOUT on seq[0] and comp rather than on the message. If your watchdog does not record that, that is the first thing to fix, and it is the subject of the Watchdog card on the Debug page.

Cascades. The first error in file order is the first error in simulation time only if you have one log. With several agents logging to sidecars, take the earliest by simulation time across all of them. Then prefer the frame closest to the root: an assertion failure on the interface beats a scoreboard mismatch two thousand cycles later, the same way ReBucket weights the top frame over the frames below it. If the assertion and the mismatch always occur together, that is a merge, not a rule.

The academic DV work is about pushing evidence closer to the root than any log message can get. The Vennsa paper built signatures from the excitation and propagation paths reported by a root-cause-analysis engine. Poulos and Veneris at ITC 2014 represented each failure as a feature vector over SAT-derived suspect sets and toggle-frequency windows, then clustered, reporting 89 percent binning accuracy and 47 percent fewer misplaced failures than message-based scripts. VCDiag in 2025 classifies failures from compressed VCD waveforms and names the top three suspect modules with over 94 percent accuracy. None of that is in reach for a nightly script today, but all of it slots into the same place: as additional fields on the Level 2 record, consumed by the same rules file. Build the pipeline so that a richer signature is a new column, not a new system.

Statistics as a debugging tool

The WER team's mantra was "data not decibels". The point of buckets is not the list; it is what you can compute over the list once the buckets are stable.

  • Pareto. WER found that a small number of buckets account for most reports. ReBucket's data for one product: 87 percent of buckets held 20 percent of hits, and 13 percent of buckets held 80 percent. Your regression will look the same. Fix the top three buckets and the tail will rise into view, which is the correct order of work.
  • New versus known. A fingerprint that appears today and did not appear yesterday is a regression. A fingerprint that appears every day is a known issue. Those two lists are the morning report, and they are a set difference between two files.
  • Bucket age and trend. First seen, last seen, hits per night. A bucket whose count is falling after a fix went in confirms the fix. A bucket whose count is rising after a fix went in is a different bug wearing the same fingerprint, which is an expanding-rule request.
  • One-hit wonders. WER reported about 10 percent of buckets with exactly one report. Do not waive them by count. A one-hit bucket that is new is a corner case the constraint solver found once; keep the seed and re-run it before deciding. A one-hit bucket that is old and never recurred is where waivers belong.
  • Owner routing. A rules file can carry an owner per fingerprint prefix, and that is enough to start. One DVCon paper trained a random forest on historical signature-to-owner assignments and reported near-perfect owner prediction for known signatures, weaker on unseen ones. The fingerprint history you accumulate is exactly that training set, so the routing table becomes a model when the table stops scaling.

The pipeline, end to end

flowchart LR
  A["Simulation
failure_stack_catcher"] -->|"one JSONL record per fail
Level 1 label"| B["Sidecar files
regression/*.jsonl"] B --> C["normalise()
mask time, seed, hex, ids"] C --> D["Level 2 evidence
last kinds, DUT state, phase"] D --> E["rules.yaml
fingerprint(), merges"] E --> F["Bucket table
count, first/last seen, owner"] F --> G["Diff vs yesterday
new / known / fixed"] G --> H["Routing
owner per bucket, one issue per bucket"]

The triage script joins the pieces. It reads every sidecar, normalises, computes the fingerprint under the current rule set, and writes a bucket table plus a diff against the previous night.

import glob, json, os, collections

def load_records(pattern):
    for path in glob.glob(pattern):
        with open(path) as f:
            for line in f:
                rec = json.loads(line)
                if rec.get("sev") in ("ERROR", "FATAL") and rec.get("first_error"):
                    rec["template"] = normalise(rec["msg"])
                    yield rec

def bucket(records, cfg):
    table = collections.defaultdict(lambda: {"hits": 0, "seeds": [], "tests": set()})
    for rec in records:
        digest, fp, title, ver = fingerprint(rec, cfg)
        b = table[digest]
        b.update(fp=fp, title=title, rules_version=ver)
        b["hits"] += 1
        if len(b["seeds"]) < 3:
            b["seeds"].append(rec["seed"])           # representative repro seeds
        b["tests"].add(rec["test"])
    return table

def diff(today, yesterday):
    new   = [k for k in today if k not in yesterday]
    fixed = [k for k in yesterday if k not in today]
    known = [k for k in today if k in yesterday]
    return new, known, fixed

cfg = yaml.safe_load(open("rules.yaml"))
today = bucket(load_records("regression/*.jsonl"), cfg)
yesterday = json.load(open("buckets/latest.json")) if os.path.exists("buckets/latest.json") else {}
new, known, fixed = diff(today, yesterday)

for k in sorted(today, key=lambda k: -today[k]["hits"]):
    b = today[k]
    flag = "NEW" if k in new else "   "
    print(f"{flag} {k} {b['hits']:5d}  {b['title']:32s}  seeds={b['seeds']}")

On a real nightly the output looks like this. The numbers are from a 190-failure run on an AXI subsystem with the rules file above at version 7.

NEW 3e9a71c0    84  SCB_MISMATCH read-path            seeds=[8812, 1207, 5590]
    a1f9c2e0    41  Byte-enable bug (merged)          seeds=[311, 9920, 4471]
    7702bd14    27  Hang in axi_burst_write_seq       seeds=[1002, 1003, 1044]
    c8d0e5f2    18  PCIe VIP CRC error                seeds=[42, 77, 5610]
    12ab8c93     9  AXI_RESP_DECERR lane 3            seeds=[6001, 6002, 6100]
    5f4e2a17     5  RAL_PREDICT_FAIL run              seeds=[913, 2211, 7370]
NEW 90c1d3ab     3  SVA a_wready_within_16            seeds=[4488, 4489, 4490]
    e77b0146     2  SCB_MISMATCH write-path           seeds=[1500, 1501]
NEW 0b6f9d55     1  UVM_FATAL cfg missing in agent[*] seeds=[2]

190 failures, 9 buckets, 3 new, 0 fixed since 2026-09-15

Read the top line the way WER would. A new bucket with 84 hits on the first night it appears is the regression; it gets the first engineer of the day. The merged byte-enable bucket is known and has an owner. The one-hit fatal with seed 2 is a config bug in a new agent instance, obvious from the title, five minutes to fix. Three people, three buckets, and nobody opened a log to decide that.

Measuring whether your buckets are honest

ReBucket was evaluated with an F-measure over hand-labelled crash data, and you can do the same at a scale that fits a Friday afternoon. Take fifty failures from one regression and have the people who fixed them write down the real bug id next to each. Then compute two numbers over your buckets:

  • Purity: for each bucket, the fraction of its failures that belong to its majority bug. Low purity means under-split; you need an expanding rule.
  • Inverse purity: for each real bug, the fraction of its failures that landed in its majority bucket. Low inverse purity means over-split; you need a condensing rule or a merge.

Keep the fifty labelled failures as a regression test for the rules file. Every rule change re-runs against them, and a change that raises purity while dropping inverse purity gets the same review a testbench change would. ReBucket reported an F-measure of about 0.88 on Microsoft's data with a learned similarity model; a hand-written rules file on a testbench you control should reach that on the second iteration, because you have something Microsoft never had: the ability to add a field to the record at the source.

Quick reference

IdeaWhere it comes fromIn the testbench
One bug per bucket, one bucket per bugWindows Error ReportingName every rule change as expanding or condensing
Group by stack, not message; top frames weigh moreSentry, Rollbar, ReBucketFailure stack: id, component, phase, sequence chain, test; test name last
In-app versus library framesSentry, ReBucket immune functionsCut paths at the VIP or uvm_pkg boundary
Strip data, keep codesRollbarMask time, seed, hex, ids, indices; keep state and opcode names and single digits
Template mining for prose logsDraindrain3 over UVM_ERROR lines when no sidecar exists
Label cheaply, classify with more evidenceWER two-phase bucketingLevel 1 in the report catcher, Level 2 in the triage script
Rules, titles, merges as dataSentry and Rollbar rule syntaxrules.yaml with version, in the testbench repo
Hangs bucket on the wait chainWER hang_wait_chainFingerprint timeouts on blocked sequence and component
Pareto, new versus known, one-hit wondersWER statisticsBucket diff every night; re-run new one-hit seeds before waiving
Purity and inverse purityReBucket evaluationFifty labelled failures as the rules file's regression test

Further reading

The Observability Triad for UVM: Logs, Metrics, and Traces Your Testbench Already Emits

The Structured Logging post turned the testbench log from prose into a queryable JSONL stream — one signal, made sharp. But a log is only one of three kinds of evidence your testbench produces on every run. Software engineering names the full set the observability triad: logs, metrics, and traces. Your UVM environment already computes all three — it counts transactions, it tracks queue depths, it threads items from sequencer to scoreboard. It just never writes two of them down. This post maps the triad onto the testbench you already have, teaches the debug motion that uses all three signals together, and builds the missing piece: a metrics collector that turns numbers you print once into curves you can query.

The Problem: One-Signal Debugging

The regression test dies at the six-hour wall-clock limit. The last line of a 400 MB log reads UVM_ERROR ... [WATCHDOG] test timeout, and it tells you nothing except that the simulation was still alive and getting nowhere when the clock ran out. It is 2 AM, the nightly is red, and the debug session begins the only way it can: grep -n ERROR run.log to find where the wheels came off, then tail -100000 run.log | less, reading backward from the corpse. Every DV engineer knows this loop. You scroll up through tens of thousands of lines looking for the moment the run turned, reconstructing a six-hour story from the one page of it that happened to print near the end.

Here is what stings about that loop: the testbench was never short on evidence. It counted every transaction that crossed every interface. It watched the depth of every scoreboard queue, every sequencer arbitration queue, every analysis FIFO. It timed every request from launch to completion. Three different kinds of evidence — how much, how deep, how long — were computed continuously, on every clock, for the entire six hours. And exactly one of them was ever written down. That one was prose: UVM_INFO, UVM_ERROR, a human sentence at a time, the weakest and least structured signal the run produced. The numbers that would actually locate the failure were calculated and thrown away.

So the log sits in front of you at 2 AM, and the questions that matter are precisely the ones it cannot answer:

  • Was throughput already sagging in the minutes before the watchdog fired, or did the whole world stop at once? A gradual slope and a cliff are different bugs — one is backpressure, the other is a deadlock — and the log flattens both into the same silence.
  • When did a scoreboard queue start growing without draining, and which queue? A depth that climbs and never recovers is the fingerprint of a dropped completion, but a depth is a number over time, and the log kept no number and no time.
  • Did this run diverge from the fifty that passed before the failure signature ever appeared? For hours this run was statistically indistinguishable from the passing population — then it was not; the log gives you the moment of death but not the moment of departure.
  • Where did the six hours actually go — which phase, which interface, which wait? Wall-clock vanished into some blocking call that never returned, and a flat stream of timestamped sentences will not tell you which one held the whole test hostage.

Read those four again and notice what they have in common: not one is a question about messages. Every one asks about a quantity over time — a rate, a depth, a divergence, a duration — or about the path a single transaction took through the environment. Those are metric questions and trace questions, and you are trying to answer them by grepping a log because the log is the only artifact the run left behind. It is the wrong instrument, aimed at the wrong signal, and no amount of grep skill fixes an instrument that never recorded the measurement.

And every one of those numbers was computed during the run. Throughput was implicit in the transaction count the monitor already kept. Queue depth was a field the scoreboard read on every push and pop. Per-request latency was a subtraction the driver could have done in its sleep. The values existed, live, in the testbench's own memory — and then the objects were garbage-collected and the values went with them. Nothing was missing at runtime. Everything was missing at debug time, because nothing was persisted.

The fix, then, is not a sharper grep and not a bigger tail. It is not even a better logging macro. The fix is to stop discarding two-thirds of the evidence — to collect the other two signals and write them down beside the log, in a form you can query after the run is dead. Software engineering has a name for the discipline of keeping all three, and a well-worn theory of what each one is for. That is where we go next.

The Triad: Three Signals from the SRE Canon

The discipline is observability, rooted in Google's Site Reliability Engineering book and named as three signals — the "three pillars" — by the observability community that built on it. Map them straight onto a testbench and the abstraction stops being abstract.

Logs are discrete, timestamped events — high detail, high volume, one line per thing that happened. They answer what happened here: this transaction, on this interface, at this nanosecond, carried this payload. Your UVM_INFO stream is already this signal, and the Structured Logging post already sharpened it.

Metrics are numeric time series — a name, a value, a timestamp, repeated on a cadence. They are cheap to store and trivial to compare across runs, and they answer how is it trending: is throughput sagging, is that queue growing, is p99 latency drifting from last night's. A metric is the number the log threw away.

Traces are causally linked event chains that share an ID — the story of one transaction as it moves through the environment. They answer where did this request go and where did it stall: sequencer to driver to DUT to monitor to scoreboard, with a timestamp at every hop, so a six-hour disappearance resolves into the one edge that took six hours.

Metrics carry more structure than a bare number, and the Prometheus type system is the vocabulary for it — because the type dictates the query. A counter is monotonically increasing; you never read it raw, you query its rate — transactions issued per microsecond is the derivative of a counter that only ever climbs. A gauge is an instantaneous level that moves both ways; you plot it as-is, because scoreboard queue depth right now is the answer. A histogram buckets a distribution; you query its percentiles, because per-transaction latency is meaningless as a mean and everything as a p50 and p99. Counters get deltas, gauges get raw plots, histograms get percentiles — pick the wrong type and the right query becomes impossible.

Which numbers do you actually collect? Two methodologies answer that, and both return in the next section's inventory. Tom Wilkie's RED method covers every request stream: Rate (how many requests per second), Errors (how many failed), Duration (the latency distribution). Point RED at a UVM interface and it becomes transactions issued, mismatches flagged, and completion latency — a counter, a counter, and a histogram, one of each.

Brendan Gregg's USE method covers every resource instead: Utilization (how busy), Saturation (how much queued work is waiting), Errors. Point USE at a scoreboard or a sequencer arbiter and it becomes occupancy, queue depth, and dropped items — gauge, gauge, counter. RED watches the flow; USE watches the thing the flow runs through, and between them they name most of what a testbench should measure.

One honest caveat, from Honeycomb's Charity Majors: the "three pillars" framing oversells the split. Logs, metrics, and traces are storage formats, not three separate kinds of truth — a single stream of wide, richly attributed events can derive all three, and treating them as three siloed backends is how you end up unable to pivot from one to another mid-debug. That critique lands even harder in DV, and in our favor. Simulation is deterministic and closed-world, and the JSONL sidecar from the Structured Logging post is already that single wide-event stream — one file, every record tagged. We are not building three backends. This post teaches three lenses on the one stream you already have.

The Mapping: What Each Signal Already Is in UVM

The theory is portable, but the payoff is specific: point each of the three signals at a UVM testbench and it lands on something already running. You are not bolting observability onto your environment. You are naming machinery that is already there and writing down two-thirds of what it produces. Here is the whole mapping on one page — the SWE incarnation, where it already lives in your testbench, and the one thing still missing.

SignalSWE incarnationAlready in your testbenchThe missing piece
Logsstdout, Splunk, ELKuvm_report_server, `uvm_infoStructure — solved in the Structured Logging post
MetricsPrometheus, DatadogCoverage %, scoreboard depths, txn counters, sim-rateNobody samples them over time — built below
TracesOpenTelemetry, Jaegerseq_item lifecycle: sequencer → driver → monitor → scoreboardA propagated ID — the next post

Read the table top to bottom and the shape of the post appears: one row is done, one row is built below, one row is teased. Take them in order.

Logs — already sharp

Logs are the row with no work left. The Structured Logging post already took the uvm_report_server stream — the same `uvm_info/`uvm_error messages you have written a thousand times — and turned it from prose into a JSONL sidecar: one JSON object per event, written to a file beside the run. The contract is five fixed fields on every record — run to name the run, ts for the simulation timestamp in nanoseconds, sev for severity, id for the report ID, and comp for the full component path — plus a msg string and whatever domain fields the event carries (an address, a transaction id, a burst length). That is the entire log signal, already structured, already queryable, already shipping. Nothing in this post revisits how to produce it; the depth lives in that post and this section defers to it wholesale.

What this post adds is the other two signals, and it writes them to that same file. A metric is not a new sidecar or a second backend — it is another line in the JSONL stream, distinguished by a single kind field. A metric record carries kind set to metric; a log record carries no kind at all, so the absence of the field is the tag that says "this is a log." One file, two record types, one query surface. That decision — one wide-event stream, not three siloed stores — is exactly the Honeycomb critique from the last section, honored by construction. The log row is closed; the metrics row is where the code starts.

Metrics — the numbers you already print once

Metrics are the meat of the mapping, because this is the signal your testbench computes most eagerly and persists least. The two methodologies from the last section tell you precisely which numbers to keep, and they partition the testbench cleanly: RED treats every interface agent as a service carrying a request stream, and USE treats every testbench structure as a resource that stream flows through.

Point RED at an interface agent. Rate is the transaction count the monitor already increments on every completed item — a counter, queried as issued-per-microsecond. Errors is the scoreboard's mismatch count, the counter that matters most because a single nonzero increment is a failing run — the error signal, already tallied on every compare. Duration is per-transaction latency from launch to completion, a histogram, and it is the leading indicator: latency starts drifting in the buckets long before the throughput counter visibly sags, so the duration signal warns you a run is sick while the rate signal still looks healthy. Rate, errors, duration — a counter, a counter, a histogram, one per interface, all three already living in fields the agent updates every clock.

Point USE at a structure. Utilization is sequencer occupancy — the fraction of time the arbiter is handing out items rather than idle — a gauge. Saturation is queue depth: the scoreboard's pending-compare queue, the analysis FIFO, the response queue, read on every push and pop. That gauge is the one that predicts death — a depth that climbs and never drains is the fingerprint of a dropped completion, visible as a rising curve minutes before the watchdog fires. Errors is the uvm_error count, the same tally the report server keeps. Occupancy, depth, error count — gauge, gauge, counter, one set per resource.

Two more numbers belong to the run as a whole rather than to any single interface or structure. Functional coverage — $get_coverage(), or per-covergroup get_coverage() — is a gauge climbing toward 100, the metric that says whether the run is doing new work or spinning. And sim-rate, the ratio of Δsim-time to Δwall-clock, paired with the monitor's txn-rate counter, is what tells apart two failures a log renders identically. The clock generator does not know or care that the DUT is stuck, so a protocol-level hang leaves sim-rate normal — clocks toggling, sim-time advancing — while the txn-rate counter goes flat, nothing completing. A collapsed sim-rate is a different animal: the simulator itself is struggling — an X-storm, runaway event activity, host memory thrashing — and the DUT is incidental to it. Neither is visible from the message stream; both are one subtraction away from numbers the run already has.

Here is the observation the whole section turns on. Every metric named above — every rate, depth, occupancy, latency, coverage percentage — already exists in your testbench today, as an instantaneous value printed exactly once, in the final report, in the report_phase, as the run winds down. Your scoreboard already prints "12,481 transactions compared, 0 mismatches, max queue depth 34." A metric is that identical number, sampled on an interval instead of totalled at the end. The difference between "prints a total when it dies" and "emits a curve you can replay" is not new measurement — the measurement is done — it is a sampling loop and a writer, roughly 150 lines, built in Build Your Own: The Metrics Collector further down. The values are already in memory. The only missing act is writing them down on a cadence.

Traces — the chain that already exists

Traces are the row this post only teases, because the causal chain is already there in your environment, just implicit. A seq_item is born in the sequence, arbitrated by the sequencer, driven onto the bus by the driver, observed by the monitor, and compared in the scoreboard — sequencer → driver → monitor → scoreboard, a fixed lifecycle every transaction walks. That path is a trace in everything but name. What it lacks is a single propagated ID that survives every hop, so that all the records belonging to one transaction can be pulled out of the stream and laid end to end. The Structured Logging post already threads a txn_id field through its records for exactly this reason; a trace is what you get when that id becomes the spine of a query.

Give the chain a propagated ID and the six-hour-disappearance question from the top of the post becomes answerable directly: follow one transaction across all four components, measure the wall-clock or sim-time spent on each hop, and the one edge that swallowed the run resolves out of the noise. That is latency-per-hop, causally attributed, no grepping required. Building it — the ID scheme, the propagation through the sequence layering, the span-per-hop model — is more than a subsection can hold. Trace IDs and Context Propagation gets its own post; it is the next card in the Foundations row. For now the mapping is enough: the trace, like the metric, is a signal your testbench already emits and simply never writes down.

The Pivot: How the Three Work Together

The three signals are not three dashboards you check in parallel — they are one motion you run in sequence, a loop that walks a metric anomaly down to a root cause: a curve bends, that names a time window; the window plus a txn_id grep names the actor that stalled; the actor's log slice names what it said when it stalled; and if the root cause is not yet in hand, the new fact points you back at a different curve. Software engineers run this loop every time a pager goes off — dashboard to trace to log, narrowing at every hop — and the mapping in the last section is what makes it run on a testbench. Each signal answers a question the previous one could not, and hands the next one a filter.

flowchart LR
    M["📈 METRICS
which curve bent, and when?"] -->|"narrows to a time window"| T["🔗 TRACES
which transaction stalled, where?"] T -->|"identifies the actor"| L["📋 LOGS
what exactly did it say?"] L -->|"new hypothesis"| M

Watch it work on a real failure — a response-starvation hang, the kind that grep renders as six hours of silence and one dying gasp. Here is what the three signals recorded, laid on a single timeline:

Sim timeSignalEvidence
28 µshistogramp99 response latency starts creeping: 180 ns → 900 ns
32 µsgaugescoreboard expected-queue depth begins monotonic growth
36 µscountertxn completion rate flatlines at zero
40 µslog[WATCHDOG] test timeout — the only line grep would have found

Read the first column and the whole argument of this post is in the gap between the top row and the bottom one. The failure began at 28 µs. The symptom — the only line the log ever wrote — landed at 40 µs. Twelve microseconds separate the moment the run turned from the moment it announced it turned, and in a message-only workflow those twelve microseconds do not exist. Grep-from-the-end starts at 40 µs, at the watchdog line, and reads backward through prose that was written while everything already looked fine on the surface, hunting for a transition that left no distinctive sentence because the failure was numeric, not textual. That is the search that takes hours: you are reconstructing a slope from a stream that only ever recorded events, scrolling up through thousands of UVM_INFO lines that each say a normal thing happened.

The metrics reader starts somewhere else entirely. The p99 latency histogram — the leading indicator from the mapping, the signal that drifts in the buckets before the rate counter visibly sags — bent at 28 µs, and it bent on the response path specifically. So before opening a single log line, the reader has two things grep never gives you: a named suspect, the response stream, and a bounded window, roughly 28 to 40 µs. The saturation gauge confirms and localizes it — the scoreboard's expected-queue depth starts its monotonic climb at 32 µs, the fingerprint of completions that are launched and never returned, and it names the scoreboard partner as the structure filling up. Two queries in, the shape of the bug is already "the responder stopped feeding the scoreboard," and the window has tightened to six microseconds.

Now the pivot. The metrics have named the where and the when; they cannot say why, because a curve is a shape and not a sentence. So you cross from the metric stream to the log stream through the one field that spans both — the timestamp — and slice the JSONL sidecar to exactly the window the curves drew:

jq 'select(.kind != "metric" and .ts > 30000 and .ts < 36000)' run.sidecar.jsonl  # ts is in ns — this is the 30–36 µs window

That is not six hours of log. It is the ~20 lines that fall inside the bent window, and read in order they tell the story the watchdog line could not: the request stream keeps going — req after req issued, monitored, accepted, the outbound side perfectly healthy — while the response stream thins, one returned completion, then a longer gap, then another, then nothing. Requests in, responses trailing off to zero, the scoreboard queue swelling behind them. Twenty lines, read once, and the mechanism is plain: the responder model has stopped producing responses while the driver keeps issuing requests. One more look at the responder's own log lines in that slice and the cause is explicit — its available-credit count reached zero and never recovered. A credit leak in the responder model: it decremented on each request but missed the return path on some completion, bled its credits to nothing, and stalled. The DUT was fine. The testbench starved itself.

Count the work. Same bug, two workflows. Log-only: open a 400 MB file at its tail, grep for the error, then read backward for an hour or more reconstructing a numeric failure from a textual record that never named it, guessing at the window because nothing marks it. Triad: one histogram query names the suspect and the window, one gauge query names the partner and confirms the mechanism, one jq slice reads the twenty lines that explain it. Three queries against one file, a few minutes, and the pivot from which curve bent to what exactly did it say is a single shared ts. The evidence was always there; the difference is that two-thirds of it got written down, and that the writing-down turns a backward scroll into a forward filter.

This is Dave Agans' third rule of debugging, "Quit Thinking and Look," with instruments instead of vibes. Agans' warning is that engineers burn hours theorizing about what might be wrong instead of observing what is wrong — and the honest reason DV engineers theorize is that looking used to mean scrolling a log, which is slow and rarely conclusive, so guessing felt faster. The triad closes that gap. Looking is now three cheap queries that converge on a six-microsecond window, and only then do you open the waveform. The waveform is still the ground truth — that is where a DV debug ends — but a waveform is six hours wide and you cannot stare at all of it. The triad's entire job is to tell you which six microseconds to open. Metrics say when, traces say which transaction, logs say what it said, and the viewer, aimed at last, says why in silicon.

Build Your Own: The Metrics Collector

The mapping section promised the missing piece was small — a sampling loop and a writer, roughly 150 lines, standing between "prints a total when it dies" and "emits a curve you can replay." Here it is, and the line count holds. But the size is the least interesting thing about it. What matters is that the design obeys three rules borrowed straight from the software side of the house, and every one of them is load-bearing. The collector must be passive: it reads the testbench, it never steers it — read-only, zero-time, no objections raised, and critically no RNG touches, because a stray $urandom inside a probe perturbs the calling thread's random state and silently breaks seed-stability of the very run you are trying to observe — read-only with respect to the testbench; a probe may keep its own bookkeeping, a delta baseline or a histogram window, but never touches the object it observes. It must be pluggable: adding a fifth metric is one new class and zero edits to the collector, or the abstraction has failed. And it must be composable: every sample lands in the same JSONL sidecar as the log records from the Structured Logging post, one wide-event stream, with the kind field doing the discrimination. Passivity keeps the instrument from changing the experiment; pluggability keeps the instrument from rotting; composability keeps you on one query surface. Hold those three in mind — every design decision below is one of them, made concrete.

Start with the contract. The whole design inverts a dependency: the collector must not know what a throughput counter or a latency histogram is, only that it can be asked its name, its type, and its current value. SystemVerilog expresses that with an interface class — a pure interface, no implementation, no state — which is the language's answer to dependency inversion.

// The contract every metric source signs. `interface class` = a pure
// interface with no implementation — SV's answer to dependency inversion.
interface class metrics_probe;
  pure virtual function string name();         // e.g. "axi_wr_throughput"
  pure virtual function string mtype();        // "counter" | "gauge" | "histogram"
  pure virtual function string sample_json();  // value as a JSON fragment; MUST NOT touch testbench state
endclass

Three methods, and the third does the real work. sample_json() returns a JSON fragment, not a fixed scalar, and that is deliberate: a gauge can hand back "14" while a histogram hands back "{\"p50\":40,\"p95\":180,\"p99\":900,\"n\":512}", and the collector splices either one into the record without caring which it got. One contract, three metric types, no if (counter) ... else if (histogram) anywhere in the collector. If the phrase "every source signs a contract and the collector iterates over the signatories" sounds familiar, it should — this is the Observer pattern from the patterns series, where the subject holds a list of observers it knows only through their interface. The collector is the subject; the probes are the observers; the interface class is the abstraction that lets the two sides evolve independently.

Now the subject itself.

class tb_metrics_collector extends uvm_component;
  `uvm_component_utils(tb_metrics_collector)

  protected metrics_probe m_probes[$];
  protected time          m_interval = 1us;   // override: +metrics_interval_ns=<n>
  protected int           m_fd;
  protected string        m_run_id;

  function new(string name, uvm_component parent);
    super.new(name, parent);
  endfunction

  function void build_phase(uvm_phase phase);
    int unsigned iv_ns;
    if ($value$plusargs("metrics_interval_ns=%d", iv_ns)) m_interval = iv_ns * 1ns;
    if (!uvm_config_db#(string)::get(this, "", "run_id", m_run_id)) m_run_id = "0";
    m_fd = $fopen({get_full_name(), ".metrics.jsonl"}, "a");
  endfunction

  function void register_probe(metrics_probe p);
    m_probes.push_back(p);
  endfunction

  // Passive: no objection raised. The test ends when the test ends;
  // the collector just stops being scheduled.
  task run_phase(uvm_phase phase);
    forever begin
      #(m_interval);
      sample_all();
    end
  endtask

  protected function void sample_all();
    foreach (m_probes[i])
      $fdisplay(m_fd, $sformatf(
        "{\"run\":\"%s\",\"ts\":%0d,\"kind\":\"metric\",\"name\":\"%s\",\"mtype\":\"%s\",\"value\":%s}",
        m_run_id, $time, m_probes[i].name(), m_probes[i].mtype(),
        m_probes[i].sample_json()));
    $fflush(m_fd);
  endfunction

  // extract_phase is the phase DESIGNED for pulling data out of the
  // testbench — the last complete sample lands even on a clean finish.
  function void extract_phase(uvm_phase phase);
    sample_all();
    $fflush(m_fd);
  endfunction

  // Every sample is flushed as it lands (sample_all), so a $fatal or a grid
  // kill still leaves the last complete window on disk. final_phase just
  // closes the handle on a normal shutdown.
  function void final_phase(uvm_phase phase);
    $fclose(m_fd);
  endfunction
endclass

A few things earn their place. The interval is not hardcoded: build_phase reads +metrics_interval_ns=<n> off the command line and falls back to the config_db run_id so the record can be correlated across runs — sample cadence is a run-time knob, not a recompile. The file opens in append mode ("a"), which is what lets these metric lines coexist with log lines in one growing stream. In the format string, note %0d on $time rather than a stringified timestamp: it keeps the ts field a bare JSON number, so jq 'select(.ts > 30000)' — the exact query from the pivot section — compares numerically instead of choking on quotes. And look hard at run_phase: it raises no objection. That single absence is the passivity rule written in code. An objection would keep the test alive to finish sampling, which means the instrument would be extending the experiment — precisely the sin the passive rule forbids. Instead the collector rides extract_phase, the UVM phase designed for pulling data out of the testbench, so the last complete sample lands even on a clean finish, and because sample_all flushes every sample as it lands, even a $fatal or a wall-clock kill — deaths no UVM phase ever sees — still leaves the last complete window on disk; final_phase merely closes the handle on a normal shutdown. The test ends when the test ends; the collector just stops being scheduled. One last note for a single-sidecar setup: swap the $fopen for a handle shared with the report server from the Structured Logging post, and the metric lines and log lines pour into the identical stream — same file, kind discriminates.

That is the whole engine. Everything from here is a probe, and each probe is one small class that signs the contract. First, a counter — per-agent write throughput, delta-sampled.

// Counter: the monitor's analysis port increments; the collector's sample
// reads the delta since last sample. Rate = delta / interval.
class axi_throughput_probe extends uvm_subscriber #(axi_txn) implements metrics_probe;
  `uvm_component_utils(axi_throughput_probe)
  protected int unsigned m_count, m_last;

  function new(string name, uvm_component parent);
    super.new(name, parent);
  endfunction

  function void write(axi_txn t);
    m_count++;                        // hot path: one increment, nothing else
  endfunction

  virtual function string name();  return "axi_wr_throughput"; endfunction
  virtual function string mtype(); return "counter";           endfunction
  virtual function string sample_json();
    int unsigned delta = m_count - m_last;
    m_last = m_count;
    return $sformatf("%0d", delta);
  endfunction
endclass

The probe subscribes to the monitor's analysis port and its write does exactly one thing on the hot path — increment — because a subscriber that does real work per transaction taxes every run whether or not anyone reads the metric. The type discipline from the theory section shows up here: a counter is never emitted raw, so sample_json() returns the delta since the last sample and stashes the new baseline, handing the querier a per-interval count it can divide into a rate. One subtlety worth calling out for anyone reaching for their linter: metrics_probe::name() does not collide with uvm_object::get_name(). The contract method was named name() precisely so it sits alongside the UVM identity method rather than overriding it — the probe answers to both, and they mean different things.

Now a gauge, and here the implements keyword shows its power. The scoreboard does not need a separate probe object watching it from outside; the scoreboard is the probe. It already knows its own queue depth, so it signs the contract directly.

// Gauge: the scoreboard IS the probe. SV's `implements` keyword lets a
// component sign the metrics contract without inheriting anything new.
class axi_scoreboard extends uvm_scoreboard implements metrics_probe;
  `uvm_component_utils(axi_scoreboard)
  protected axi_txn m_expected[$];
  // ... normal scoreboard duties elided ...

  virtual function string name();  return "sb_depth"; endfunction
  virtual function string mtype(); return "gauge";    endfunction
  virtual function string sample_json();
    return $sformatf("%0d", m_expected.size());  // read-only: .size(), no pop
  endfunction
endclass

A gauge is an instantaneous level, so sample_json() just reports m_expected.size() as-is — no delta, no reset. And note what it does not do: it calls .size(), never .pop_front(). Reading the depth must not consume the queue, or the instrument would be altering the thing it measures. This is the same saturation gauge that climbed at 32 µs in the pivot timeline — the fingerprint of completions launched and never returned — now emitted on a cadence instead of printed once as "max queue depth 34."

The third type is the histogram, and it is the one the pivot section leaned on hardest.

// Histogram: monitors report request->response latency; the probe sorts
// the current window at sample time and emits percentiles, then resets.
class axi_latency_probe extends uvm_subscriber #(axi_txn) implements metrics_probe;
  `uvm_component_utils(axi_latency_probe)
  protected time m_window[$];

  function new(string name, uvm_component parent);
    super.new(name, parent);
  endfunction

  function void write(axi_txn t);
    m_window.push_back(t.rsp_time - t.req_time);
  endfunction

  virtual function string name();  return "axi_rd_latency"; endfunction
  virtual function string mtype(); return "histogram";      endfunction
  virtual function string sample_json();
    time sorted[$] = m_window; string s;
    if (sorted.size() == 0) return "{\"n\":0}";
    sorted.sort();
    s = $sformatf("{\"p50\":%0d,\"p95\":%0d,\"p99\":%0d,\"n\":%0d}",
                  sorted[(sorted.size()-1)*50/100],   // (n-1)*p/100: always in range
                  sorted[(sorted.size()-1)*95/100],
                  sorted[(sorted.size()-1)*99/100],
                  sorted.size());
    m_window.delete();
    return s;
  endfunction
endclass

The write path stays cheap — one subtraction pushed onto a window — and the expensive step, the sort, happens only at sample time when someone is actually going to read the number. The percentile index (n-1)*p/100 is chosen so it always lands inside the array for any non-empty window, which is why the size-zero case returns early with {"n":0}. Then the window is cleared, so each sample describes the latencies in that interval rather than the run to date. This is the leading indicator made concrete: p99 on the response path is the signal that bent at 28 µs in the pivot timeline — a full 8 µs before the throughput counter flatlined at 36. The histogram warns you a run is sick while the counter still looks healthy, and now that warning is a queryable curve.

Three metric types, three probes, and every one of them lived entirely inside the HDL. The fourth reaches for something SystemVerilog does not have. Sim-rate is Δsim-time over Δwall-clock, and the language has no wall-clock — $time is simulated nanoseconds, and $system() shells out for a status code, not a value you can read back. This is the moment to open the software toolbox and pull out DPI-C. Three lines of C give the testbench a monotonic wall clock.

// SystemVerilog has no wall-clock. $system() returns a status, not a value.
// Three lines of C fix that:
import "DPI-C" function longint tb_epoch_ms();
// tb_epoch.c — compile into the sim with your tool's DPI flow
#include <time.h>
long long tb_epoch_ms(void) {
  struct timespec ts;
  clock_gettime(CLOCK_MONOTONIC, &ts);
  return (long long)ts.tv_sec * 1000 + ts.tv_nsec / 1000000;
}

With a wall clock imported, the probe is another gauge.

// Gauge: sim-nanoseconds advanced per wall-millisecond. Read WITH the
// txn-rate counter: txn rate flat while sim_rate stays normal = protocol
// hang (clocks still toggling, nothing completing); sim_rate collapsed =
// the simulator itself is struggling (X-storm, runaway event activity,
// host memory thrash).
class sim_rate_probe implements metrics_probe;
  protected time    m_last_sim;
  protected longint m_last_wall;

  virtual function string name();  return "sim_rate"; endfunction
  virtual function string mtype(); return "gauge";    endfunction
  virtual function string sample_json();
    time    now_sim  = $time;
    longint now_wall = tb_epoch_ms();
    real    rate     = (now_wall == m_last_wall) ? 0.0
                     : real'(now_sim - m_last_sim) / real'(now_wall - m_last_wall);
    m_last_sim  = now_sim;
    m_last_wall = now_wall;
    return $sformatf("%.1f", rate);
  endfunction
endclass

The comment carries the diagnostic payload, and it is the one from the mapping section, unchanged: read sim-rate with the txn-rate counter and the pair separates two failures a log renders identically. Txn rate flat while sim-rate stays normal is a protocol hang — clocks toggling, sim-time advancing, nothing completing. Sim-rate collapsed is a different animal — the simulator itself is struggling under an X-storm, runaway event activity, or host memory thrash, and the DUT is incidental. One more structural note: sim_rate_probe extends nothing. It is a plain class, not a uvm_component, because a probe needs only to sign the contract — the collector holds it by its metrics_probe handle and never asks it to be anything more. That is the dependency inversion paying off; the collector's list does not care what class of object it holds. And with four probes in hand, the fifth writes itself: a coverage_pct gauge — $get_coverage() wrapped in the same three-method shape — is the reader's first exercise; it is six lines.

Which leaves the wiring, and the wiring is where the whole design either earns its keep or does not.

function void tb_env::connect_phase(uvm_phase phase);
  super.connect_phase(phase);
  axi_agent.mon.ap.connect(m_wr_tput.analysis_export);
  axi_agent.mon.ap.connect(m_rd_lat.analysis_export);
  m_metrics.register_probe(m_wr_tput);   // counter
  m_metrics.register_probe(m_rd_lat);    // histogram
  m_metrics.register_probe(m_sb);        // gauge — the scoreboard itself
  m_metrics.register_probe(m_sim_rate);  // gauge — plain class
endfunction

The subscribers hook onto the monitor's analysis port, and then all four probes register with the collector through the same one-line register_probe call — a counter, a histogram, and two gauges, one of which is the scoreboard measuring itself and one of which is a plain non-component class, handled identically because the collector sees only the contract. This is the pluggability rule collecting its debt: adding metric #5, that coverage_pct gauge, means writing one class and adding one register_probe line here — it touches zero lines inside tb_metrics_collector, because the collector was never told what any of its probes are. That is dependency inversion earning its keep, and it is why the whole collector fits in 150 lines while staying open to every metric your testbench will ever want to emit.

Querying the Signals

Everything in the pivot section ran on one query primitive: jq, pointed at the JSONL sidecar, filtering on kind and ts. The filenames below are illustrative, not prescriptive — a metrics sidecar and a log sidecar, shown here as two files. Section 5's collector writes wherever you point its $fopen: one shared stream discriminated by kind, or two files that cat concatenates into one — the queries below run unchanged either way. That is not a coincidence to gloss over — it is the entire payoff of writing one wide-event stream instead of three siloed backends. You do not need a metrics database, a trace viewer, or a query language with a learning curve. You need four one-liners, and you already have three of them memorized from the sections above. Here they are as a toolkit, each doing one job, none of them needing anything more exotic than the jq already on your workstation.

The first turns a single metric into a plottable series — pick a name, drop everything else, keep timestamp and value:

# One metric as CSV — feed to any plotter
jq -r 'select(.kind=="metric" and .name=="sb_depth") | [.ts, .value] | @csv' run.metrics.jsonl

Against the fixture built for this section — sb_depth sampled every microsecond, climbing from 32 µs — this line reads out 0,1, 1000,0, 2000,1, in order, ready for any plotter that takes CSV.

The second answers a narrower question — not the whole series, but the one instant that matters: where does the counter last read nonzero, the sample just before the flatline the pivot section built its whole case around.

# Hang window: last sample where throughput was still nonzero
jq -r 'select(.name=="axi_wr_throughput" and (.value|tonumber) > 0) | .ts' run.metrics.jsonl | tail -1

On the same fixture — throughput steady at 8 through 35 µs, zero from 36 µs on — this returns 35000: one number, ns, and it is the right edge of the window you need to slice next.

The third is the one line that reaches straight into a histogram's structure rather than a scalar, because the percentile is stored inline on every sample and needs no recomputation at query time:

# p99 latency curve — histograms carry their percentiles inline
jq -r 'select(.name=="axi_rd_latency") | [.ts, .value.p99] | @csv' run.metrics.jsonl

Run against the fixture, the p99 column sits flat at 180 through 28 µs and then climbs — 230, 280, 330 — exactly the drift the pivot's histogram query depended on, read straight off .value.p99 with no post-processing.

The fourth is the pivot itself, restated as a reusable line rather than a one-off: metrics named the window, now slice the log stream — the records with no kind field — to that same span.

# The pivot: metrics found the window, now slice the LOG stream to it
jq 'select(.kind != "metric" and .ts > 30000 and .ts < 36000)' run.sidecar.jsonl

Four lines, one file format, no new tool. Piped to a CSV, a metric series is one step from a picture, and a picture is what makes a bent curve visible at a glance instead of buried in a column of numbers. Feed two series through gnuplot and the throughput drop and the depth climb from the pivot's timeline land on one plot, one axis each:

set datafile separator ","
set y2tics
plot "tput.csv"  using 1:2 with lines title "wr throughput",      \
     "depth.csv" using 1:2 with lines title "sb depth" axes x1y2

One PNG per failing run, generated in the regression triage script — the curve is the first debug step.

The real payoff is not one run's plot — it is fifty of them. Run the same jq → CSV pipeline over every passing run from the last week, overlay all fifty curves on the failing run's curve on one axis, same metric, same color for the passing band, one line standing out in the failing run's color. Fifty healthy throughput traces sit in a tight band; the failing run's trace rides inside that band for hours and then peels away from it, and the instant it peels away is not a guess — it is a pixel you can point at. That instant is the debug starting line. It is earlier than the watchdog line by however many microseconds the pivot section found, and it is exactly the moment worth asking "what changed here" about, because the answer to that question is the bug.

Fifty passing runs and one failing run, overlaid, is one query away — but ten thousand runs a night, with dozens of distinct failure signatures tangled together, is a different problem than eyeballing one divergent line on one plot. That is where this series goes next: the Triage category picks up from here, with an Error Fingerprinting card asking what pattern a failure leaves behind when you cannot look at each of ten thousand plots by hand. The signals are captured; querying them at that scale is the next post's problem, not this one's.

Quick Reference

Everything above compresses to one table, one checklist, and a reading list. Bookmark this section; it is the part you come back to at 2 AM instead of re-reading the whole post.

SignalUVM sourceCollection costFirst question it answers
Logsuvm_report_server → JSONL sidecarDone — previous postWhat exactly happened at time T?
Metrics (counter)analysis-port subscriber, delta-sampled~20 lines/probeWas it still making progress?
Metrics (gauge)scoreboard/queue .size() via implements~10 lines/probeWhat was filling up, and when did it start?
Metrics (histogram)latency window, percentiles at sample time~30 lines/probeWas it getting slower before it stopped?
Tracestxn_id threading (next post)One UUID + disciplineWhere did THIS transaction stall?

Adopt the triad in the order that pays off fastest, not the order this post presented it:

  1. Drop in tb_metrics_collector with the sim_rate probe — zero testbench coupling, and your regression can already tell a hung DUT from a thrashing simulator.
  2. Register scoreboard-depth (implements metrics_probe, ~10 lines) and one per-agent throughput probe.
  3. Add the jq → gnuplot step to the fail-triage script: every failing run gets its curves rendered before a human looks at it.

Each step stands alone — step 1 alone already answers the "is the simulator or the DUT stuck" question from the mapping section, without touching a single scoreboard or agent class.

Further reading:

  • Google, Site Reliability Engineering — the chapter on monitoring distributed systems, the monitoring foundations under the three-signal framing this post maps onto UVM.
  • Peter Bourgon — Metrics, Tracing, and Logging — the post that named the three pillars.
  • Charity Majors (Honeycomb) — the "three pillars" critique: logs, metrics, and traces are storage formats, not separate truths.
  • Brendan Gregg — the USE method (Utilization, Saturation, Errors), for naming what to measure on any resource.
  • Tom Wilkie — the RED method (Rate, Errors, Duration), for naming what to measure on any request stream.
  • Prometheus documentation — metric types (counter, gauge, histogram), the vocabulary this post borrows wholesale.
  • OpenTelemetry — semantic conventions, the closest thing the industry has to a standard schema for traces.
  • David Agans, Debugging — Rule 3, "Quit Thinking and Look," the rule the pivot section runs on.

That closes the loop this post opened. The Structured Logging post gave the testbench its first sharp signal; this one added the two it was throwing away. Next in Foundations: Trace IDs & Context Propagation — the UUID that turns your log stream into traces, and the piece this post kept teasing and deferring. Until then, everything here — logs, metrics, and the queries that connect them — lives on the Debug hub, alongside the rest of the triage toolkit this series is building one card at a time.

State Pattern in UVM: Behavior That Changes With State, Without the Case Explosion

The Template Method post fixed the flow in one place and let each test override only the steps. Now invert it: the methods stay fixed, and what changes is what each call does — because the answer depends on where the object sits in its lifecycle. The pattern that hands each state its own class — so each state's behavior lives in one place instead of being scattered across walls of branching — is State. The question it answers comes straight from a MESI scoreboard: a snoop read arrives for an address your reference model tracks. If the line is Modified, it must supply data and downgrade; if it's Invalid, it must do nothing at all. Same event, opposite behavior — and that's one event out of four, two cells of a sixteen-cell matrix. How many case arms before the branching owns you?

The Problem: The MESI Matrix in a Case Statement

Your scoreboard's reference model tracks the MESI state of every cache line the test has touched: core reads and writes arrive from the CPU-side monitor, snoops arrive from the bus monitor, and on each event the model must predict the line's next state and whatever bus traffic the transition owes. Four states, four event kinds — the protocol is a sixteen-cell matrix, and the spec draws the legal transitions as a single diagram:

stateDiagram-v2
    [*] --> I
    I --> S : core read (fill)
    I --> M : core write (RFO)
    S --> M : core write (upgrade)
    E --> M : core write (silent)
    E --> S : snoop read
    M --> S : snoop read (supply data)
    M --> I : snoop invalidate (writeback)
    S --> I : snoop invalidate
    E --> I : snoop invalidate

(One simplification, then we move on: a read miss fills to Shared in this model; real protocols fill to Exclusive when no other cache holds the line.)

The natural first implementation maps each event to a function and each function to a case over the state. Here are two of the four — including the snoop-read handler that answers the question from the top of the post, one arm supplying data and downgrading, another doing nothing at all:

// Inside the reference model — one of FOUR functions shaped exactly like this:
function void on_core_write(cache_line_t line);
  case (line.state)
    INVALID:   begin expect_rfo(line);    line.state = MODIFIED; end
    SHARED:    begin expect_upgrade(line); line.state = MODIFIED; end
    EXCLUSIVE: line.state = MODIFIED;     // silent upgrade — no bus traffic
    MODIFIED:  ;                          // already ours
  endcase
endfunction

function void on_snoop_read(cache_line_t line);
  case (line.state)
    MODIFIED:  begin expect_supply(line); line.state = SHARED; end
    EXCLUSIVE: line.state = SHARED;
    SHARED:    ;                          // clean copy elsewhere — memory answers
    INVALID:   ;                          // not ours... but is silence checked, or assumed?
    default:   `uvm_error("MESI", "unreachable")  // this handler has a default. The one above doesn't.
  endcase
endfunction

// ... on_core_read and on_snoop_invalidate: two MORE case (line.state) blocks ...

Sixteen cells of protocol, sliced event-first into four functions. It compiles, it passes the smoke test, and three problems are already loaded and waiting:

  • MESI → MOESI is a shotgun edit across every event handler — and it misses one. The project moves to a protocol with the Owned state, so OWNED needs an arm in all four functions: what an Owned line does on a core write, on a core read, on each snoop — and it doesn't only add arms, it changes an existing one, because on_snoop_read's MODIFIED arm now goes M → O, supplying data with no memory writeback, instead of M → S. You open the handlers and add the arm. You add it to three. The fourth — say on_snoop_invalidate, the one you skipped because "invalidate is invalidate" — has no OWNED arm and no default, so an Owned line sails through the case untouched, stays Owned in the model after the bus said otherwise, and surfaces as a scoreboard mismatch forty thousand cycles downstream with no arrow pointing back at the arm you never wrote.
  • Illegal-transition handling is inconsistent, because the policy exists nowhere in particular. Read the two handlers again: on_snoop_read carries a default arm that fires a uvm_error; on_core_write carries none, so anything unexpected falls through silently. Neither author was careless — there was simply no single place where "what happens on an impossible transition" got decided, so each handler decided alone. The same goes for the INVALID arm of the snoop handler: whether the model checks that a line it doesn't own stays off the bus, or merely assumes it, depends on nothing but the habits of whoever wrote that arm — the question was never decided, so the code answers it differently wherever it comes up.
  • You cannot read the protocol back out of the code. The spec writes per-state bullet lists — "a Modified line supplies data on a snoop read, writes back on an invalidate, absorbs core writes silently." To review the code against that one bullet list you collect one arm from each of four case statements, scattered across the file. The spec's unit of meaning is the state; the code's unit of organization is the event — and every review, every debug session, every new hire reads against that grain.

What you want is each state owning its own behavior — its arm from every one of those four case statements gathered into one class, so the code reads in the spec's units. That's the State pattern.

Gang of Four: The State Pattern

Before the cache line gets its classes, take the Gang of Four version on its own, so the structure stands clear of any MESI detail.

"Allow an object to alter its behavior when its internal state changes. The object will appear to change its class."

— Design Patterns: Elements of Reusable Object-Oriented Software (Gamma et al., 1994)

The second sentence is the strange one, and the whole pattern hides inside it. An object cannot change its class — but it can hold a handle to another object that can be swapped, and route all of its behavior through that handle. Swap the object behind the handle and, from the outside, the context appears to have become something else.

classDiagram
    class Context {
        -state: State
        +request()
    }
    class State {
        <<abstract>>
        +handle(ctx)* State
    }
    class ConcreteStateA {
        +handle(ctx) State
    }
    class ConcreteStateB {
        +handle(ctx) State
    }
    Context o--> State : delegates to
    State <|-- ConcreteStateA
    State <|-- ConcreteStateB
    note for Context "request() {
state = state.handle(this);
}"

Read the note on Context — that one line is the entire mechanism. The context owns a state handle and does no deciding of its own: every event that arrives, it delegates to whatever state it currently holds. Each concrete state implements handle() with the behavior that state owes and returns the successor state — the return value is the transition. (GoF's own sample plumbs this differently — the state reaches back and calls setState() on the context — but returning the successor is the same decision with the mutation moved into one line, and it's the form this post uses throughout.) The context never consults a transition table, never branches on an enum; it stores whatever came back and the next event lands on the new object.

Successor selection lives inside the states themselves: ConcreteStateA knows which state follows it, because that knowledge is the protocol — exactly the per-state bullet list the spec writes and the case statements of §1 shredded across four functions.

Now put that diagram next to Strategy's — the most-confused sibling pair in the catalog — and they are nearly identical: a context, an abstract interface, concrete subclasses behind it, behavior swapped by swapping an object. The one tell in our diagram — handle() returning a State — is precisely the difference the table below makes explicit: who drives the swap.

AspectStateStrategy
Who selects the objectThe current state, by returning its successorThe client, once, at injection
Objects aware of each otherYes — states name their successorsNo — strategies are independent
Changes during runConstantly, driven by eventsRarely or never
IntentBehavior follows a lifecycleInterchangeable algorithms

Rule of thumb: if the object picks its own successor, it's State; if the testbench picks it once at config time, it's Strategy.

UVM Implementation: State Machines Everywhere, State Pattern Nowhere

Every post in this series anchors the pattern in a place where UVM itself ships it. This one arrives with a twist: UVM is full of state machines — and not one uses the State pattern. That is not a gap in the library; it is the boundary of the pattern, and reading why the framework stayed on the enum side teaches when you should cross over. Two sightings make the case.

The first is the largest state machine in the framework: phasing. Every uvm_phase node carries its own lifecycle, walked through UVM_PHASE_DORMANT → SCHEDULED → … → EXECUTING → … → DONE — nine states along the normal walk, with SYNCING, STARTED, READY_TO_END, ENDED, and CLEANUP filling the gaps. Multiply that by every phase node in every domain and an ordinary simulation runs dozens of live state machines. And the API for all of them? get_state() returns the enum, wait_for_state() blocks until it matches. (One trap the glimpse below sidesteps: call them on the schedule node the domain's find() resolves — the uvm_main_phase::get() imp singleton never changes state.)

The second sighting is one you have called yourself, perhaps without filing it as a state machine: uvm_sequence_base tracks every sequence through its own enum lifecycle — CREATED → PRE_BODY → BODY → FINISHED (or STOPPED), with more waypoints than this sketch shows. Again the public surface is a query and a wait: get_sequence_state() and wait_for_sequence_state().

// UVM's state machines are everywhere — and they're all enums:
uvm_phase main_ph = uvm_domain::get_uvm_domain().find(uvm_main_phase::get());
main_ph.wait_for_state(UVM_PHASE_EXECUTING);          // phase lifecycle

my_seq.wait_for_sequence_state(UVM_FINISHED);          // sequence lifecycle

// Consumers only QUERY and WAIT — never branch behavior per state.
// That's exactly when an enum is the right tool — and exactly what
// stops being true inside your MESI reference model.

Now the honest part. Both are plain enums, and that is the correct choice, not a missed refactor. What consumers do with these states: ask which one is current, and block until a particular one arrives. Nothing more. The framework's own traversal machinery does branch on these states internally — but that branching is a fixed walk over a frozen state set, not a roster of event handlers where each state answers differently, not new states arriving release over release, not a sixteen-cell matrix. When every use is a comparison, a case never gets the chance to explode. The boundary in one sentence: enum + case when the state is only ever read — compared, waited on, stamped into a field; State pattern when behavior varies per state. UVM's lifecycles sit on the first side; the §1 reference model — where a snoop read means "supply data and downgrade" in one state and "do nothing" in another — sits on the second.

That sentence unpacks into a checklist for any state machine you meet (it returns as a table in §7):

  • Stay with enum + case when the arms are one-liners — assignments, comparisons, waits — and the state set is frozen. A case over four states that only stamps a field is not debt; it is the simplest correct tool.
  • Reach for State when the arms grow behavior — calls, side effects, expected-traffic bookkeeping; when the state×event matrix grows in both dimensions (MESI → MOESI adds a row, a new snoop type a column, every cell an edit site); or when illegal-transition policy must be uniform and auditable instead of re-decided per handler, the way §1's two case statements decided it two different ways.

The verdict on UVM is clean — query-and-wait consumers, frozen state sets — so this section's GoF mapping cannot point at the framework. It points at what §4 builds; the framework supplies only the base classes and the object discipline, the roles yours to fill:

GoF RoleYour Class (built in §4)
Contextcache_line (per-line model object inside the scoreboard)
State (abstract)mesi_state virtual base class
handle() event methodshandle_core_read/core_write/snoop_read/snoop_invalidate(cache_line line)
ConcreteStatemodified_state, exclusive_state, shared_state, invalid_state
Transitioneach handler returns the next mesi_state

UVM's state machines never needed the State pattern. Your reference model is about to.

Build Your Own: A MESI Cache Line That Knows Its State

Time to fill §3's table. The plan is §2's mechanics applied to §1's matrix: a mesi_state virtual base class declaring one handler per event, four concrete states each collecting its arm from each of the four case statements, and a cache_line context that holds the current state handle and delegates everything. Each handler returns the successor — the return value is the transition — and the sixteen-cell matrix re-slices itself state-first, into the spec's per-state bullet lists.

The Abstract State: mesi_state

Start with the base class, because it is where §1's second pain point goes to die. Four virtual handlers, one per event kind, each taking the cache_line context and returning the next state — and every base implementation fires a uvm_error, so the question §1's two handlers answered two different ways ("what happens on an impossible transition?") is now decided in exactly one place. A concrete state that doesn't override a handler has declared that event illegal, and the inherited error is uniform, auditable, and impossible to forget: the missing-default arm and the silent fall-through can no longer be written, because not writing anything now fires it.

typedef class cache_line;     // states and context name each other — forward-declare

virtual class mesi_state extends uvm_object;
  // ... constructor ...
  // Every handler returns the NEXT state. The base class IS the
  // illegal-transition policy: one place, uniform, auditable.
  virtual function mesi_state handle_core_read(cache_line line);
    `uvm_error("MESI", $sformatf("illegal core read in %s", get_name()))
    return this;
  endfunction
  virtual function mesi_state handle_core_write(cache_line line);
    `uvm_error("MESI", $sformatf("illegal core write in %s", get_name()))
    return this;
  endfunction
  virtual function mesi_state handle_snoop_read(cache_line line);
    `uvm_error("MESI", $sformatf("illegal snoop read in %s", get_name()))
    return this;
  endfunction
  virtual function mesi_state handle_snoop_invalidate(cache_line line);
    `uvm_error("MESI", $sformatf("illegal snoop invalidate in %s", get_name()))
    return this;
  endfunction
endclass

One consequence worth pausing on: this base class also forces the question §1's INVALID snoop arm left to authorial habit. An invalid_state that wants a snoop read to be a legal no-op must say so — override the handler, return this — turning silence-as-policy into a line of code you can point at in review.

The Concrete States: the Spec's Bullet Lists, as Classes

Now the states themselves. Here are Invalid and Shared in full — read either top to bottom and you are reading the spec's bullet list for that state, all four events in one place. (The ::get() calls are singleton accessors — one shared instance per state instead of a new() per transition, §5's business; for now read shared_state::get() as "the Shared state object.")

typedef class shared_state;    // states name their successors — mutual references
typedef class modified_state;  // need forward declarations within one file

class invalid_state extends mesi_state;
  // ... constructor, ::get() singleton — see Scaling Up ...
  virtual function mesi_state handle_core_read(cache_line line);
    line.expect_fill();                  // bookkeeping lives in the context
    return shared_state::get();          // the return value IS the transition
  endfunction
  virtual function mesi_state handle_core_write(cache_line line);
    line.expect_rfo();                   // write miss — read-for-ownership
    return modified_state::get();
  endfunction
  virtual function mesi_state handle_snoop_read(cache_line line);
    return this;                         // not ours — an EXPLICIT legal no-op
  endfunction
  virtual function mesi_state handle_snoop_invalidate(cache_line line);
    return this;
  endfunction
endclass

class shared_state extends mesi_state;
  // ... constructor, ::get() singleton ...
  virtual function mesi_state handle_core_read(cache_line line);
    return this;                         // read hit — stay Shared
  endfunction
  virtual function mesi_state handle_core_write(cache_line line);
    line.expect_upgrade();               // must invalidate other sharers
    return modified_state::get();
  endfunction
  virtual function mesi_state handle_snoop_read(cache_line line);
    return this;                         // someone else supplies; memory answers
  endfunction
  virtual function mesi_state handle_snoop_invalidate(cache_line line);
    return invalid_state::get();
  endfunction
endclass

// exclusive_state: read hit stays E; write silently upgrades to M;
//                  snoop read downgrades to S; snoop invalidate to I.
// modified_state:  hits stay M; snoop read supplies data, downgrades to S;
//                  snoop invalidate writes back, goes to I.

The question from the top of the post is now two adjacent overrides of the same handler: modified_state supplies data and returns shared_state::get(); invalid_state returns this. Same event, opposite behavior — one method name, two classes.

(The typedef class lines are not decoration. Invalid names Shared and Modified as successors, Shared names Invalid back — states knowing their successors is the pattern, per §2 — and mutual references in one file require forward declarations or the compile fails on the first shared_state::get().)

Note what invalid_state does not contain: a single branch. Neither does shared_state. The case statements are gone — not moved, gone — because an object that exists only while the line is Invalid never needs to ask what state the line is in. And note the division of labor, §2's diagram made concrete: the state selects the successor and what traffic the transition owes; the context — line.expect_fill(), line.expect_rfo(), line.expect_upgrade() — does the bookkeeping. (Per §1's simplification, expect_rfo is the write-miss path, expect_fill the read-miss path.)

The Context: cache_line and the Scoreboard That Stops Deciding

The context is deliberately boring. It owns the per-line data — tag, data, expected-traffic bookkeeping — plus one protected handle to the current state, an accessor, and a mutator:

class cache_line extends uvm_object;
  bit [TAG_W-1:0]  tag;
  bit [63:0]       data;
  protected mesi_state m_state;
  // ... constructor: m_state = invalid_state::get(); ...
  function mesi_state state();  return m_state;  endfunction
  function void set_state(mesi_state next);      // ONE chokepoint — §5 instruments it
    m_state = next;
  endfunction
  // expect_fill / expect_rfo / expect_upgrade / expect_supply / expect_writeback:
  // bookkeeping the scoreboard checks against observed bus traffic
endclass

// The scoreboard's analysis export — the case MATRIX is gone;
// what's left is one-level event decode:
function void write_snoop(snoop_txn t);
  cache_line line = get_line(t.addr);  // addr → line map — grows in §5
  case (t.kind)
    SNOOP_READ: line.set_state(line.state().handle_snoop_read(line));
    SNOOP_INV:  line.set_state(line.state().handle_snoop_invalidate(line));
  endcase
endfunction

Look at what survived in the scoreboard's write(): one case — over the event kind, not the state. That branching is irreducible — something has to map a transaction field to a method call — but it is one level deep, never grows a state dimension, and contains no behavior. The state×event matrix — sixteen cells across four functions, headed for twenty-five once MOESI adds the row and the next snoop type the column — is not in this function. Each line is §2's note on Context, flattened: the scoreboard's write plays request(), composing delegate-and-store through the public accessor pair, keeping set_state the one visible chokepoint §5 will instrument. The line "appears to change its class" with every set_state, and the next snoop lands on whatever object the last transition chose.

Now collect the payoff against §1's three pain points, in order. MESI → MOESI stops being a shotgun edit. owned_state is a new class — one file, one bullet list, written next to the spec's O-state paragraph — plus a one-line change in modified_state::handle_snoop_read to return it. No fourth handler to forget, because there are no handlers to visit: a transition you fail to write isn't a silent fall-through forty thousand cycles from a mismatch, it's the base class's uvm_error on the first Owned-line event you didn't think about. Existing states closed to modification, the state set open to extension — Open/Closed, the same inversion every post in this series keeps arriving at. Illegal-transition policy lives in one place. Decided once, in mesi_state, inherited uniformly, with legal no-ops spelled out as explicit return this overrides you can audit. And you can finally read the protocol back out of the code. The spec's unit of meaning is the state; now it is the code's too — reviewing modified_state against the spec's Modified bullet list is a side-by-side read, not a scavenger hunt through four case statements.

Pitfalls

Three traps, each quietly rebuilding the problem you just removed:

  • State objects accumulating per-line data. The first time a handler needs the line's tag, the temptation is to give the state a tag field. Do it and the design is broken: there are thousands of cache lines and (if §5's singletons are doing their job) exactly one shared_state, so any per-line fact stored in a state is shared by every line in that state — a corruption bug that only fires when two lines occupy the state at once. States must stay stateless; everything per-line lives in the context, which is why every handler takes cache_line line as an argument.
  • new()-ing a state per transition. shared_state nxt = new("s"); return nxt; works — and allocates an object per event, at analysis-port rates, for objects that carry no data and never differ: millions of snoops, millions of identical throwaway allocations. The ::get() singletons elided above are the fix, and building them is where §5 picks up.
  • Transition side effects scattered into the states. When invalid_state needs a fill expected, it calls line.expect_fill() — it does not reach into the scoreboard's queues itself. Hold that line: selection belongs to the states (which successor, which traffic is owed), mutation belongs to the context (do the bookkeeping, own the data, swap the handle). Blur it — states pushing into scoreboard queues, or worse, calling set_state on the side while also returning a successor — and transitions happen in two places, the chokepoint stops being a chokepoint, and the transition logging §5 hangs on it sees only half the story.

Two states in full, two elided, one boring context — and a scoreboard that no longer decides anything. What's left is what this section deferred: those ::get() calls, and what one chokepoint and one shared instance per state buy when the model grows from one line to a coherent reference model.

Scaling Up: From One Line to a Coherent Reference Model

A real coherence test does not touch one line. The addr → line map that §4's scoreboard waved past fills with every address the test visits — a hundred thousand cache_line objects is an ordinary long run — and events arrive at analysis-port rates across all of them. That is the scale at which §4's three IOUs come due: that stateless states make a shared instance safe, that ::get() retires the per-transition new(), and that a single set_state chokepoint would earn its keep. Cash them in.

Four Objects for a Hundred Thousand Lines

Here is the ::get() every handler in §4 was already calling — the standard lazy-init Singleton: a local static handle, constructed on first request, returned ever after.

class shared_state extends mesi_state;
  local static shared_state m_inst;
  static function shared_state get();
    if (m_inst == null) m_inst = new("shared_state");
    return m_inst;                       // 100k lines, ONE Shared object
  endfunction
  virtual function mesi_kind_e kind();  return MESI_S;  endfunction
  // ... handlers as in Build Your Own ...
endclass

(That kind() override is new — hold it; it belongs to the chokepoint, not the singleton.)

Now pay off both §4 pitfalls at once. Why it's safe: states stay stateless — every per-line fact (tag, data, expected traffic) lives in the cache_line context, arriving as a handler argument. A shared_state carries nothing that distinguishes one line from another, so a hundred thousand lines pointing m_state at the same object cannot interfere — there is nothing to interfere with. Why it's worth it: new()-per-transition would mean millions of identical throwaway allocations; with ::get(), the reference model holds exactly four state objects for the whole simulation, and a transition is a handle copy — no allocation, no garbage, no churn.

One tease before moving on: "many contexts sharing a few stateless instances" is not just a Singleton trick — it has its own name in the Gang of Four catalog, Flyweight, and its own post. Here it rides in on Singleton mechanics, but the intent is Flyweight's.

The Transition Chokepoint

§4's third pitfall insisted that every transition flow through set_state — states return successors, never swapping the handle themselves. Here is what that buys: because exactly one line of code changes a cache line's state, instrumenting it instruments every transition in the run:

// set_state is now declared extern in cache_line — this out-of-class
// body replaces §4's inline one. Two new members support it:
// cur_event: mesi_event_e, stamped by the scoreboard (line.cur_event = SNOOP_READ;)
//            before delegating — forget the stamp and every log line below
//            misreports its event;
// trans_cg:  covergroup with function sample(mesi_kind_e s, mesi_event_e e,
//            mesi_kind_e n) — (state,event)→next_state cross, new()'d in
//            the constructor.
function void cache_line::set_state(mesi_state next);
  if (next != m_state) begin
    `uvm_info("MESI/TRANS",
      $sformatf("line=%0h %s -> %s on %s",
                tag, m_state.get_name(), next.get_name(), cur_event.name()),
      UVM_HIGH)
    trans_cg.sample(m_state.kind(), cur_event, next.kind());
  end
  m_state = next;
endfunction

Two payloads, one guard. The uvm_info is a structured log line — fixed ID, fixed old → new on event shape — so grep 'MESI/TRANS' over a failing run replays the model's entire decision history, and "every line that left Modified on a snoop" is one filter, not a debug session. The trans_cg.sample call gives you transition coverage for free: a (state, event) → next_state cross sampled on every real transition, with no per-test, per-sequence, or per-state sampling code anywhere — there is nowhere else a transition can happen. And the next != m_state guard keeps the return this no-ops out of both — log and coverage: self-loops are legal but not transitions, and at scoreboard rates mostly noise — a comparison only meaningful because return this and ::get() hand back the same object every time, so the singletons quietly underwrite the guard. Whether they belong in the coverage is a separate decision the guard is making for you — a read hit in Modified is a cell of §1's matrix too — so if your plan wants the diagonal covered, sample before the guard, or cover (state, event) occupancy in a second, unguarded group.

One bridge made that sample call compile. Covergroups sample integral values, not class handles — so the mesi_state base gains pure virtual function mesi_kind_e kind(); (legal in a virtual class, and unlike §4's handlers there is no sensible default, so pure forces every concrete state to answer), and a four-value mesi_kind_e enum — MESI_M, MESI_E, MESI_S, MESI_I — reappears in the package. Yes: the enum §1 deleted is back. But read what it's for — kind() is a reporting detail, an integral shadow for sampling and logging, and nothing branches on it. The day a case (line.state().kind()) shows up, §1 has been rebuilt with extra steps. NEVER case on kind().

Reset: The Fifth Handler

§4's base class declared exactly four handlers, one per event the bus could deliver. Reset grows the roster for the first time, and it breaks the base-class rule the other four established: reset is the one event no steady state needs to customize — whatever you were, clear the line and become Invalid.

// In the mesi_state BASE — the one transition every state shares.
// The only base handler that is not an illegal-transition error:
virtual function mesi_state handle_reset(cache_line line);
  line.clear();                            // (returns invalid_state::get() —
  return invalid_state::get();             //  same forward-declaration dance as §4)
endfunction

(This model treats reset as invalidate-without-writeback — line.clear() drops Modified's dirty data on the floor, what most hardware does. A DUT that flushes on reset is one override away: modified_state alone reclaims handle_reset, calls expect_writeback(), and falls through to the base — the default-in-the-base shape was built for exactly this.)

The shape is the payoff inverted. The four event handlers put the error in the base because legality varies per state; handle_reset puts the behavior in the base because it doesn't — M, E, S, and I answer identically, so it is written once and no concrete state in THIS model mentions reset. And because the scoreboard routes reset like any other event — line.set_state(line.state().handle_reset(line)) — every reset lands in the chokepoint: the M → I on RESET lines show up in the log and the reset column fills in the covergroup, zero reset-specific instrumentation. (One caveat: a reset on an already-Invalid line is an I→I self-loop the guard clips, so the (I, RESET) cell fills only if you sample the diagonal as above.)

Worth noticing what just happened to the base class, because it will happen again: §4's handler set was four, reset makes it five, and §6 adds a sixth when transient states need an event the steady states never see. The abstract state's method list is the protocol's event vocabulary, and it grows when the vocabulary does — one declaration in the base per new event, against §1's one new arm in every case statement.

Four shared objects, one instrumented chokepoint, one universal event written once: the model now scales to a hundred thousand lines and tells you what it did. What it still cannot say is that a line is between states — a fill issued but not yet answered, an eviction in flight. Giving the in-between a class of its own is where the pattern gets interesting — and it's next.

Advanced: Transient States

The model the last five sections built carries one comfortable lie: that MESI has four states and a line always sits in exactly one. Between a core write to an I line and the fill data arriving from memory, the line is neither Invalid nor Modified — it is in flight, an "I→M pending fill," obligated to a transaction that has not completed. Add the acks a real protocol collects before it may upgrade and there are more still. This is why a production coherence model carries not four states but dozens, most of them these transients — the brief, obligated moments between the stable four. The pattern's promised payoff is that absorbing one costs almost nothing. Time to make good on the promise, and to be honest about the word "almost."

One New State, Counted Honestly

Here is "I→M pending fill" as a state object. The line has issued its read-for-ownership and is waiting; when the fill lands it commits the deferred write and becomes Modified; a snoop arriving meanwhile cannot be answered — the data has not arrived — so it stalls.

// New protocol reality: the fill takes time. One NEW class:
class im_pending_fill_state extends mesi_state;
  // ... constructor, ::get() singleton ...
  virtual function mesi_kind_e kind();  return MESI_IM;  endfunction
  virtual function mesi_state handle_fill(cache_line line);
    line.commit_pending_write();
    return modified_state::get();
  endfunction
  virtual function mesi_state handle_snoop_read(cache_line line);
    line.stall_snoop();                  // can't supply what we don't have yet
    return this;
  endfunction
  // Everything else inherits the base illegal-transition policy — free.
endclass

// One new event in the BASE (illegal everywhere by default):
virtual function mesi_state handle_fill(cache_line line);
  `uvm_error("MESI", $sformatf("illegal fill in %s", get_name()))
  return this;
endfunction

// And ONE changed return value — the transition that enters the new state:
class invalid_state extends mesi_state;
  virtual function mesi_state handle_core_write(cache_line line);
    line.expect_rfo();
    return im_pending_fill_state::get();  // was: modified_state::get()
  endfunction
  // ... rest unchanged ...
endclass

Now count the cost honestly, because the temptation is to call this "zero edits" and it is not. The transient state costs three behavioral edits and one bookkeeping touch — and naming that fourth out loud is the whole point, because "zero edits" is the story that hides it. One, a new class — im_pending_fill_state, written next to the spec's pending-fill paragraph, overriding only the two events it has an opinion about, inheriting illegal-by-default for the rest. Two, one new handler in the base: handle_fill, which mesi_state declares illegal everywhere, so every other state — M, E, S, I — gets "a fill here is a protocol violation" for free, the illegal-by-default discipline §4 established. Three — the one the "zero edits" story drops — one changed return value: invalid_state::handle_core_write now returns im_pending_fill_state::get() instead of modified_state::get(), because a write miss now parks in the pending state until the fill commits it. One line, one state's logic, not a character of M, E, or S. Four — and §5's pure-virtual kind() already forced this — the new class returns a fresh MESI_IM, so the coverage-only mesi_kind_e enum grows by one. Still nothing cases on it. That kind() override lives inside the one new class of item one, but the enum bump is a genuine touch in another file, so it earns its own number. Not zero — four touches, every one local, named, and pointable in review.

Set that against §1's case matrix. A new state there is a new enum value and therefore a new arm in every event handler — not "add an arm" but audit one: you reopen all four (now five, six) case statements and decide, for each, what a pending-fill line does, including the handlers where the answer is "illegal" and the only way to say so is a default you might skip. The pattern turns "audit every arm of every case statement" into "write one class and redirect one line" — localized extension versus distributed audit, the whole reason the pattern earns its keep, sharpest precisely here where states arrive by the dozen.

This is the sixth handler. §4 declared four — one per event the bus delivers. §5's reset made five. handle_fill is the sixth, the kind §5 foreshadowed: an event the steady states never see — meaningless to Invalid, Modified, or Shared, since only a line caught mid-transition is waiting for one. So the base declares it illegal everywhere and exactly one transient state overrides it, the inverse of handle_reset: reset belongs to every state and lives in the base; fill belongs to almost none and lives in the one state that owns the in-between.

The Hazard: Class Explosion, and the Honest Exit Sign

Every post in this series ends on the pattern's signature failure mode, and State's is written on the wall the moment you say "dozens of transient states." §1's case matrix obscured the protocol horizontally — one state's behavior shredded across four functions. A model with dozens of transient states reintroduces the same disease rotated ninety degrees: the protocol is now obscured vertically — dozens of small classes, most three lines long, each a single obligated moment, the shape lost in the file list. You can read any one state perfectly and still not see the machine. The cure for horizontal scatter became vertical scatter.

Three things keep it readable, and the third matters most because it tells you when to stop.

  • Name the two tiers apart. Stable and transient states are different animals — one is where a line rests, the other where a line is briefly obligated. A convention like mesi_* for the stable four and mesi_t_* for every transient (mesi_t_im_pending_fill) makes the file list itself legible: the four states that hold the protocol's shape stand apart from the dozens that thread between them.
  • Keep the map in one place. The stable-state diagram and the full transient list belong in a single doc block at the top of the package, not discovered class by class. The classes hold the behavior; one block holds the shape, so a reviewer sees the whole machine before reading any one cell — the vertical answer to §1's horizontal scavenger hunt.
  • And the exit sign, stated plainly. Watch what the transient handlers become as they multiply. The interesting ones — im_pending_fill_state, with its stall and deferred commit — carry real behavior and belong in classes. But the more you add, the more handlers shrink to a single line: return next_state::get(), no side effect, no stall, no bookkeeping — pure transition. When most of your states do nothing but name a successor, the per-state class has stopped paying for itself: you are spending a class, a singleton, and a kind() override to encode one arrow. That is the honest exit sign. A protocol whose transients are almost all pure plumbing is asking for a table-driven FSM — the transition relation as data, a literal (state, event) → next_state table the engine walks — not a wall of three-line classes encoding the same table one method at a time. The pattern is not the destination; it is the right tool while behavior varies per state, and the discipline to leave it when behavior drains out is the same judgment §3 used to keep UVM's lifecycles on the enum side. Reach for State when states act; reach for the table when they merely point.

Seventeen patterns. Seventeen orthogonal concerns. Factory builds, Adapter translates, Bridge decouples, Decorator adds, Facade simplifies, Proxy mediates, Composite treats a leaf and a subtree the same way, Observer broadcasts, Strategy swaps the algorithm, Chain of Responsibility delegates until claimed, Command encapsulates, Template Method fixes the flow — and State lets behavior follow the lifecycle, one class per state. None replaces another.

Quick Reference

Everything in one place — the §3 mapping finalized with the transient additions, the §3 checklist delivered as the table it promised, and the traps to avoid:

GoF Role → UVM Mapping

GoF RoleYour Class / Mechanism
Contextcache_line (per-line model object inside the scoreboard)
State (abstract)mesi_state virtual base class
handle() event methodshandle_core_read/core_write/snoop_read/snoop_invalidate + handle_reset (§5) + handle_fill (§6)
ConcreteStatemodified_state, exclusive_state, shared_state, invalid_state, im_pending_fill_state
Transitioneach handler returns the next mesi_state

Every concrete state also implements kind() — not a GoF role, but the mandatory coverage-only integral shadow from §5, the fourth touch every new state owes.

Enum + case vs State Pattern: The Decision Checklist

SignalVerdict
Case arms are one-liners (assign, compare, wait)enum + case
State set is frozenenum + case
Consumers only query and wait (get_state/wait_for_state)enum + case
Arms grow behavior (calls, side effects, bookkeeping)State pattern
The state×event matrix grows in both dimensionsState pattern
Illegal-transition policy must be uniform and auditableState pattern

Common Mistakes

MistakeFix
State objects holding per-line dataKeep states stateless; pass the context into every handler
new() per transitionShared (singleton) state instances — all mutable data in the context
Illegal transitions silently hit defaultBase-class handlers uvm_error by default; legal rows override
Transition side effects scattered across handlersOne set_state() chokepoint — logging, coverage, mutation in one place
Confusing State with StrategyStates pick their own successors at runtime; a Strategy is injected once
Rebuilding the case matrix inside one state classOne virtual method per event — never case (event) inside a state

Previous: Template Method Pattern — Define the flow once, override only the steps

Next: Coming soon