Error Fingerprinting for UVM Regressions: Bucketing 10,000 Failures the Way Sentry and Windows Do
The overnight regression finished with 10,234 failures. The structured-logging post ended with a one-line jq that collapses them into a dozen buckets by hashing three fields. That line works, and it is also where most teams stop. Then, two weeks later, the buckets start lying: a scoreboard message that hides five different bugs, an address in a message that splits one bug into four hundred buckets, a hang that lands in whichever bucket the watchdog message happens to match. Nobody trusts the dashboard, and the senior engineer is back to opening logs in tabs.
The software world hit exactly this wall twenty-five years ago and wrote down what it learned. Microsoft's Windows Error Reporting has bucketed crash reports from a billion machines since 1999. Sentry and Rollbar bucket exceptions for hundreds of thousands of applications and expose their grouping rules in public documentation. Microsoft Research published the algorithm that replaced message matching with call-stack matching. This post takes those results and rebuilds error fingerprinting for UVM regressions on top of them: what a fingerprint must be orthogonal to, what the "stack trace" of a testbench failure actually is, how to normalise a message without destroying its signal, and how to keep the rules honest as the project changes.
- Two ways a fingerprint lies
- The failure stack: what your testbench's stack trace is
- Normalising the message without losing the signal
- Label in the simulation, classify in the script
- Rules, overrides and merges
- The first error is not always the root
- Statistics as a debugging tool
- The pipeline, end to end
- Measuring whether your buckets are honest
- Quick reference
Two ways a fingerprint lies
The Windows Error Reporting team defined the goal in one sentence: a bucketing algorithm should maintain orthogonality, one bug per bucket and one bucket per bug. Every failure of a fingerprint is a failure of one half of that sentence.
- Under-split: two bugs in one bucket. Your scoreboard reports
SCB_MISMATCHfor every data error it sees. A parity bug in the write path and a byte-lane bug in the read path both surface asSCB_MISMATCHfrom the same component. Bucketed on message id and component, they are one bucket with 1,450 hits, one owner, and one very confused afternoon. - Over-split: one bug in many buckets. The same scoreboard prints the address in the message. One bug, hit by 400 seeds at 400 addresses, becomes 400 buckets of one hit each. The dashboard says "400 distinct problems" and the team stops reading it.
The Vennsa and University of Toronto paper that called failure triage "the neglected debugging problem" drew exactly these two pictures for hardware: two distinct bugs caught by the same checker, and one bug caught by different checkers because different stimulus propagated it along different paths. Their observation was that the industry's automation, where it existed, binned "purely on the error message and the owners of the failing tests", and that this was the source of both failure modes.
WER's vocabulary for fixing this is worth adopting as-is. A heuristic that adds information to the fingerprint is expanding: it increases the bucket count so that two bugs stop sharing one. A heuristic that removes information is condensing: it decreases the bucket count so that one bug stops spanning several. Their table of client-side heuristics is mostly expanding (add the module name, the offset, the exception code, the hang wait-chain root) with a few deliberate condensing ones (replace the module and offset with an in-code assert ID, because the assert ID identifies the bug better than where it fired). Every change you make to a fingerprint rule is one of these two moves, and you should know which one you are making and why.
uvm_report id instead of the message text. The equivalent of the expanding hang_wait_chain heuristic is fingerprinting a timeout by which sequence was blocked and on what, not by the watchdog's generic message.The failure stack: what your testbench's stack trace is
Sentry's default grouping does not look at the message first. It looks at the stack trace, and it distinguishes in-app frames, your code, from library and framework frames, which it ignores for grouping. Rollbar does the same: its exception fingerprint is a hash of the exception class plus the file and method names of the frames, with line numbers dropped and framework boilerplate frames removed. The 2012 ReBucket paper from Microsoft Research went further and replaced exact matching with a similarity measure: frames near the top of the stack (the crash point) weigh more than frames near the bottom, and two stacks that match the same functions at slightly different depths still count as similar. ReBucket also strips "immune functions", library code that is trusted enough to be unlikely to hold the bug.
A UVM failure has a stack trace too. It is just not in the place a software engineer would look. The frames, from top to bottom, are:
| Frame | Source | Weight | Why |
|---|---|---|---|
| Assertion or checker id | rm.get_id(), property name | Highest | Closest to the observation of the bug, the "crash point" |
| Reporting component | rm.get_report_object().get_full_name(), instance indices masked | High | Which monitor or scoreboard saw it |
| Phase | uvm_phase at the time of the error | Medium | Reset-phase failures and run-phase failures are rarely the same bug |
| Running sequence chain | get_parent_sequence() walked to the root | Medium, decaying downward | What stimulus was active; the outer virtual sequence matters less than the innermost |
| Test name | +UVM_TESTNAME | Lowest | The frame most scripts bucket on, and the least informative: one test hits many bugs and one bug hits many tests |
Notice the weight column is the inverse of common practice. Grouping by test name is grouping by the bottom frame, the main() of the failure. ReBucket's position-dependent model says that is the frame to trust least.
The in-app rule also transfers cleanly. Components in your environment are in-app. Anything under uvm_pkg and inside a purchased VIP is a library frame: keep it for context, drop it from the fingerprint. A failure reported by uvm_test_top.env.pcie_vip.dl_layer.crc_checker is fingerprinted at the VIP boundary, env.pcie_vip, plus the VIP's own error id, because the frames inside it are not yours to fix and their internal paths change with every VIP release.
Here is the collector. It runs inside a report catcher, so it sees every error at the moment it is raised, and it emits one compact record per failing simulation, on the first error only.
class failure_stack_catcher extends uvm_report_catcher;
`uvm_object_utils(failure_stack_catcher)
bit captured;
string library_roots[$] = '{"pcie_vip", "axi_vip"};
// Set from phase_started() in your base test: current_phase = phase.get_name();
static string current_phase = "build";
function new(string name = "failure_stack_catcher");
super.new(name);
endfunction
// Mask instance indices: env.axi_agent[3].monitor -> env.axi_agent[*].monitor
function string mask_indices(string path);
string out = "";
foreach (path[i]) begin
if (path[i] == "[") begin
out = {out, "[*]"};
while (i < path.len() && path[i] != "]") i++;
end else out = {out, path[i]};
end
return out;
endfunction
// First index of sub in s, or -1 (SystemVerilog has no built-in substring search)
function int index_of(string s, string sub);
for (int i = 0; i + sub.len() <= s.len(); i++)
if (s.substr(i, i + sub.len() - 1) == sub) return i;
return -1;
endfunction
// Cut the path at a library boundary so VIP internals do not enter the fingerprint
function string to_in_app(string path);
foreach (library_roots[k]) begin
int idx = index_of(path, library_roots[k]);
if (idx >= 0) return path.substr(0, idx + library_roots[k].len() - 1);
end
return path;
endfunction
function action_e catch();
uvm_report_object ro;
uvm_sequence_base seq;
string frames[$];
if (captured || get_severity() < UVM_ERROR) return THROW;
captured = 1;
ro = get_client();
frames.push_back({"id:", get_id()});
frames.push_back({"comp:", mask_indices(to_in_app(ro.get_full_name()))});
frames.push_back({"phase:", current_phase});
// Sequence chain, innermost first, stored as kinds not instances
if ($cast(seq, ro) == 0) seq = running_sequence(ro);
while (seq != null) begin
frames.push_back({"seq:", seq.get_type_name()});
seq = seq.get_parent_sequence();
end
log_failure_record(frames, get_message());
return THROW;
endfunction
endclass
Three hooks are yours. Override phase_started() in the base test with one line, failure_stack_catcher::current_phase = phase.get_name();, so the catcher knows the phase without reaching into UVM internals. running_sequence() and log_failure_record() are the other two: the first resolves the sequence currently on the sequencer that drives the reporting agent, the second appends one JSON line to the sidecar file described in the structured-logging post. The record carries the frames, the raw message, and the run header fields (seed, test, RTL hash, tool version). Everything below consumes that record.
Normalising the message without losing the signal
Rollbar publishes its message-normalisation list, and it is short: strip dates, timestamps, email addresses, IP addresses, decimal numbers, integers of two or more digits, and hex values of four or more digits, then hash what is left. The interesting part is the exception: error codes and HTTP status codes are deliberately kept, because a 404 and a 500 in the same handler are different bugs. "Strip things that look like data, keep things that look like codes" is the whole rule.
The DV version of the list:
| Mask | Pattern | Keep or strip | Reason |
|---|---|---|---|
| Simulation time | @ 12345 ns, time=... | Strip | Changes with every seed |
| Seed | seed=..., +ntb_random_seed | Strip | Run identity, not bug identity |
| Addresses and data | 0x[0-9a-f]{4,}, 'h... | Strip | The classic over-split source |
| Transaction and packet ids | id=NN, seq_id, tag= | Strip | Per-run counters |
| Instance indices | [3] in paths | Strip | Same bug across ports |
| Small integers | 0 to 9 standing alone | Keep | Usually a lane, a channel, a state: a code, not data |
| State and opcode names | IDLE, L0s, WRITE_BURST | Keep | The DV equivalent of an error code |
| Checker verdict words | expected, actual, timeout, underflow | Keep | Distinguish mismatch from protocol violation |
The one judgment call is the two-digit rule. Rollbar strips integers of two or more digits and keeps single digits. That heuristic is right for hardware too: a lane index of 3 is a code, a burst length of 256 is data. Apply the same rule and you get 90 percent of the benefit with no per-message configuration.
import re
MASKS = [
(r"@\s*\d+\s*[pnuf]?s\b", "@<T>"), # sim time
(r"\b(seed|ntb_random_seed)\s*[=:]\s*\d+", r"\1=<SEED>"),
(r"0x[0-9a-fA-F]{4,}", "<HEX>"), # addresses, data
(r"'h[0-9a-fA-F_]{4,}", "<HEX>"),
(r"\b(id|tag|seq_id|txn)\s*[=:]\s*\d+", r"\1=<ID>"),
(r"\[\d+\]", "[*]"), # instance indices
(r"\b\d{2,}\b", "<N>"), # integers of 2+ digits
]
def normalise(msg: str) -> str:
for pat, rep in MASKS:
msg = re.sub(pat, rep, msg)
return re.sub(r"\s+", " ", msg).strip()
# "SCB_MISMATCH addr=0x4000_1000 lane 3 expected 0xdead actual 0xbeef @ 128340 ns"
# -> "SCB_MISMATCH addr=<HEX> lane 3 expected <HEX> actual <HEX> @<T>"
If your logs are prose because the testbench predates structured logging, you do not have to write masks by hand. The Drain algorithm, published in 2017 and maintained as the drain3 package, mines message templates from a stream of log lines: variable positions become <*> automatically, a similarity threshold decides when a line starts a new template, and custom masks for hex and integers are applied first. Feed it your UVM_ERROR lines and the template id it returns is a usable normalised message on day one.
Label in the simulation, classify in the script
WER's most transferable design decision is that bucketing happens in two phases in two places. Labeling runs on the client, from the evidence available at the moment of the crash, and is cheap. Classifying runs on the server, once more data has arrived (symbols, memory dumps, a second report from the same bucket), and can move a report to a better bucket. Bucketing is progressive: the first answer is not the final answer.
Your simulation is the client. Your triage script is the server.
- Level 1, in the simulation: a label from immediate evidence only. Top two frames plus the normalised message:
id|component|template. It costs nothing, needs no other run, and is right often enough to route the failure. - Level 2, in the triage script: a classification that adds evidence the simulation could not see cheaply or at all. The last five transaction kinds from the ring buffer. The DUT state signature. Whether the failure was preceded by a reset or a power transition. Whether the same seed passed on yesterday's RTL. This is where the scoreboard's five hidden bugs get separated, because the read-path bug always has a
READ_BURSTin the last five kinds and the write-path bug never does.
Two records that share a Level 1 label but split at Level 2 tell you the label was under-split, and that is an expanding heuristic waiting to be promoted into the simulation-side rule. Two Level 2 buckets that a human keeps merging tell you the opposite. The levels are not just a pipeline. They are how the rule set learns.
!analyze tool grew to roughly 100,000 lines implementing about 500 bucketing heuristics, at about one new heuristic a week for a decade. That number is the honest forecast for any fingerprint rule set: it will keep changing. Store the rule-set version in every record, and never compare buckets across versions without re-bucketing the old records.Rules, overrides and merges
Sentry and Rollbar both expose the same three controls, and their syntaxes are worth copying because they were arrived at by years of users asking for the same things.
Fingerprint rules match an event and replace the fingerprint. Sentry writes them as one line each, first match wins:
error.type:DatabaseUnavailable -> system-down
error.value:"connection error: *" -> connection-error, {{ transaction }}
logger:my.package.* level:error -> error-logger, {{ logger }} title="Error from Logger {{ logger }}"
Rollbar writes them as JSON with a condition, a fingerprint template and a title, and its condition language has the operators you would expect: equality, membership, prefix and suffix, numeric comparison, regex, combined with any, all and none. Both let a rule refer back to the default fingerprint, so a rule can refine grouping instead of replacing it.
Merges are the human override for under-split rules: two issues merged so that future events land in one, with an unmerge that restores the originals. Sentry now also proposes merges from a transformer embedding of stack traces, which is where "similar enough" grouping ends up once exact matching runs out of road.
For a regression flow the same three controls fit in one YAML file next to the testbench, evaluated by the triage script in order:
version: 7
rules:
# Condensing: every VIP-internal CRC error is one bug at the VIP boundary
- match: { id: "PCIE_DL_CRC*", comp: "env.pcie_vip*" }
fingerprint: "pcie-vip-crc|{{ phase }}"
title: "PCIe VIP CRC error"
# Expanding: the scoreboard hides read-path and write-path bugs behind one id
- match: { id: "SCB_MISMATCH", last_kinds: "*READ_BURST*" }
fingerprint: "{{ default }}|read-path"
# Timeouts are grouped by what was blocked, never by the watchdog text
- match: { id: "WATCHDOG_TIMEOUT" }
fingerprint: "hang|{{ seq[0] }}|{{ comp }}"
title: "Hang in {{ seq[0] }}"
merges:
- into: "a1f9c2e0"
from: ["7b3d0e11", "c04e9a77"]
reason: "Same byte-enable bug seen by monitor and scoreboard, JIRA DV-418"
The evaluator is short. Conditions are shell-style globs against record fields, templates pull fields into the fingerprint string, and the hash is taken over the final string so that rules can be reordered without changing existing bucket ids.
import fnmatch, hashlib, json, re, yaml
def matches(rule, rec):
return all(fnmatch.fnmatch(str(rec.get(k, "")), pat)
for k, pat in rule["match"].items())
def render(template, rec, default):
def sub(m):
key = m.group(1).strip()
if key == "default":
return default
if key.startswith("seq["):
return rec["seq"][int(key[4:-1])] if rec.get("seq") else ""
return str(rec.get(key, ""))
return re.sub(r"{{\s*(.*?)\s*}}", sub, template)
def fingerprint(rec, cfg):
default = "|".join([rec["id"], rec["comp"], rec["template"]]) # Level 1 label
fp, title = default, rec["id"]
for rule in cfg["rules"]:
if matches(rule, rec):
fp = render(rule["fingerprint"], rec, default)
title = render(rule.get("title", title), rec, default)
break
digest = hashlib.sha1(fp.encode()).hexdigest()[:8]
for m in cfg.get("merges", []):
if digest in m["from"]:
digest = m["into"]
return digest, fp, title, cfg["version"]
Three properties of this design matter more than the code. Rules live in version control next to the testbench, so a rule change is reviewed like any other change. The rule-set version travels with every record. And the merge table is data, not a special case, so an override made at 2 AM survives the next rule edit.
The first error is not always the root
Every fingerprint pipeline starts by taking the first error in the log. Two situations break that assumption, and both have a software precedent.
Hangs. WER's hang heuristic does not fingerprint a hang by the fact that the UI stopped. It walks the chain of threads waiting on synchronisation objects, starting from the input thread, to find the root of the wait chain, and buckets on that. The DV equivalent: a UVM_FATAL from the watchdog is the observation, not the bug. The fingerprint should come from the watchdog's structured report of what was blocked, the sequence that never got its response, the sequencer it was waiting on, and the phase, which is why the rule above fingerprints WATCHDOG_TIMEOUT on seq[0] and comp rather than on the message. If your watchdog does not record that, that is the first thing to fix, and it is the subject of the Watchdog card on the Debug page.
Cascades. The first error in file order is the first error in simulation time only if you have one log. With several agents logging to sidecars, take the earliest by simulation time across all of them. Then prefer the frame closest to the root: an assertion failure on the interface beats a scoreboard mismatch two thousand cycles later, the same way ReBucket weights the top frame over the frames below it. If the assertion and the mismatch always occur together, that is a merge, not a rule.
The academic DV work is about pushing evidence closer to the root than any log message can get. The Vennsa paper built signatures from the excitation and propagation paths reported by a root-cause-analysis engine. Poulos and Veneris at ITC 2014 represented each failure as a feature vector over SAT-derived suspect sets and toggle-frequency windows, then clustered, reporting 89 percent binning accuracy and 47 percent fewer misplaced failures than message-based scripts. VCDiag in 2025 classifies failures from compressed VCD waveforms and names the top three suspect modules with over 94 percent accuracy. None of that is in reach for a nightly script today, but all of it slots into the same place: as additional fields on the Level 2 record, consumed by the same rules file. Build the pipeline so that a richer signature is a new column, not a new system.
Statistics as a debugging tool
The WER team's mantra was "data not decibels". The point of buckets is not the list; it is what you can compute over the list once the buckets are stable.
- Pareto. WER found that a small number of buckets account for most reports. ReBucket's data for one product: 87 percent of buckets held 20 percent of hits, and 13 percent of buckets held 80 percent. Your regression will look the same. Fix the top three buckets and the tail will rise into view, which is the correct order of work.
- New versus known. A fingerprint that appears today and did not appear yesterday is a regression. A fingerprint that appears every day is a known issue. Those two lists are the morning report, and they are a set difference between two files.
- Bucket age and trend. First seen, last seen, hits per night. A bucket whose count is falling after a fix went in confirms the fix. A bucket whose count is rising after a fix went in is a different bug wearing the same fingerprint, which is an expanding-rule request.
- One-hit wonders. WER reported about 10 percent of buckets with exactly one report. Do not waive them by count. A one-hit bucket that is new is a corner case the constraint solver found once; keep the seed and re-run it before deciding. A one-hit bucket that is old and never recurred is where waivers belong.
- Owner routing. A rules file can carry an owner per fingerprint prefix, and that is enough to start. One DVCon paper trained a random forest on historical signature-to-owner assignments and reported near-perfect owner prediction for known signatures, weaker on unseen ones. The fingerprint history you accumulate is exactly that training set, so the routing table becomes a model when the table stops scaling.
The pipeline, end to end
flowchart LR A["Simulation
failure_stack_catcher"] -->|"one JSONL record per fail
Level 1 label"| B["Sidecar files
regression/*.jsonl"] B --> C["normalise()
mask time, seed, hex, ids"] C --> D["Level 2 evidence
last kinds, DUT state, phase"] D --> E["rules.yaml
fingerprint(), merges"] E --> F["Bucket table
count, first/last seen, owner"] F --> G["Diff vs yesterday
new / known / fixed"] G --> H["Routing
owner per bucket, one issue per bucket"]
The triage script joins the pieces. It reads every sidecar, normalises, computes the fingerprint under the current rule set, and writes a bucket table plus a diff against the previous night.
import glob, json, os, collections
def load_records(pattern):
for path in glob.glob(pattern):
with open(path) as f:
for line in f:
rec = json.loads(line)
if rec.get("sev") in ("ERROR", "FATAL") and rec.get("first_error"):
rec["template"] = normalise(rec["msg"])
yield rec
def bucket(records, cfg):
table = collections.defaultdict(lambda: {"hits": 0, "seeds": [], "tests": set()})
for rec in records:
digest, fp, title, ver = fingerprint(rec, cfg)
b = table[digest]
b.update(fp=fp, title=title, rules_version=ver)
b["hits"] += 1
if len(b["seeds"]) < 3:
b["seeds"].append(rec["seed"]) # representative repro seeds
b["tests"].add(rec["test"])
return table
def diff(today, yesterday):
new = [k for k in today if k not in yesterday]
fixed = [k for k in yesterday if k not in today]
known = [k for k in today if k in yesterday]
return new, known, fixed
cfg = yaml.safe_load(open("rules.yaml"))
today = bucket(load_records("regression/*.jsonl"), cfg)
yesterday = json.load(open("buckets/latest.json")) if os.path.exists("buckets/latest.json") else {}
new, known, fixed = diff(today, yesterday)
for k in sorted(today, key=lambda k: -today[k]["hits"]):
b = today[k]
flag = "NEW" if k in new else " "
print(f"{flag} {k} {b['hits']:5d} {b['title']:32s} seeds={b['seeds']}")
On a real nightly the output looks like this. The numbers are from a 190-failure run on an AXI subsystem with the rules file above at version 7.
NEW 3e9a71c0 84 SCB_MISMATCH read-path seeds=[8812, 1207, 5590]
a1f9c2e0 41 Byte-enable bug (merged) seeds=[311, 9920, 4471]
7702bd14 27 Hang in axi_burst_write_seq seeds=[1002, 1003, 1044]
c8d0e5f2 18 PCIe VIP CRC error seeds=[42, 77, 5610]
12ab8c93 9 AXI_RESP_DECERR lane 3 seeds=[6001, 6002, 6100]
5f4e2a17 5 RAL_PREDICT_FAIL run seeds=[913, 2211, 7370]
NEW 90c1d3ab 3 SVA a_wready_within_16 seeds=[4488, 4489, 4490]
e77b0146 2 SCB_MISMATCH write-path seeds=[1500, 1501]
NEW 0b6f9d55 1 UVM_FATAL cfg missing in agent[*] seeds=[2]
190 failures, 9 buckets, 3 new, 0 fixed since 2026-09-15
Read the top line the way WER would. A new bucket with 84 hits on the first night it appears is the regression; it gets the first engineer of the day. The merged byte-enable bucket is known and has an owner. The one-hit fatal with seed 2 is a config bug in a new agent instance, obvious from the title, five minutes to fix. Three people, three buckets, and nobody opened a log to decide that.
Measuring whether your buckets are honest
ReBucket was evaluated with an F-measure over hand-labelled crash data, and you can do the same at a scale that fits a Friday afternoon. Take fifty failures from one regression and have the people who fixed them write down the real bug id next to each. Then compute two numbers over your buckets:
- Purity: for each bucket, the fraction of its failures that belong to its majority bug. Low purity means under-split; you need an expanding rule.
- Inverse purity: for each real bug, the fraction of its failures that landed in its majority bucket. Low inverse purity means over-split; you need a condensing rule or a merge.
Keep the fifty labelled failures as a regression test for the rules file. Every rule change re-runs against them, and a change that raises purity while dropping inverse purity gets the same review a testbench change would. ReBucket reported an F-measure of about 0.88 on Microsoft's data with a learned similarity model; a hand-written rules file on a testbench you control should reach that on the second iteration, because you have something Microsoft never had: the ability to add a field to the record at the source.
Quick reference
| Idea | Where it comes from | In the testbench |
|---|---|---|
| One bug per bucket, one bucket per bug | Windows Error Reporting | Name every rule change as expanding or condensing |
| Group by stack, not message; top frames weigh more | Sentry, Rollbar, ReBucket | Failure stack: id, component, phase, sequence chain, test; test name last |
| In-app versus library frames | Sentry, ReBucket immune functions | Cut paths at the VIP or uvm_pkg boundary |
| Strip data, keep codes | Rollbar | Mask time, seed, hex, ids, indices; keep state and opcode names and single digits |
| Template mining for prose logs | Drain | drain3 over UVM_ERROR lines when no sidecar exists |
| Label cheaply, classify with more evidence | WER two-phase bucketing | Level 1 in the report catcher, Level 2 in the triage script |
| Rules, titles, merges as data | Sentry and Rollbar rule syntax | rules.yaml with version, in the testbench repo |
| Hangs bucket on the wait chain | WER hang_wait_chain | Fingerprint timeouts on blocked sequence and component |
| Pareto, new versus known, one-hit wonders | WER statistics | Bucket diff every night; re-run new one-hit seeds before waiving |
| Purity and inverse purity | ReBucket evaluation | Fifty labelled failures as the rules file's regression test |
Further reading
- Glerum et al., Debugging in the (Very) Large: Ten Years of Implementation and Experience, SOSP 2009.
- Dang et al., ReBucket: A Method for Clustering Duplicate Crash Reports Based on Call Stack Similarity, ICSE 2012.
- Sentry, fingerprint rules and stack trace rules.
- Rollbar, default grouping algorithm and custom fingerprinting rules.
- He et al., Drain: An Online Log Parsing Approach with Fixed Depth Tree, ICWS 2017 (Drain3 implementation).
- Safarpour et al., Failure Triage: The Neglected Debugging Problem, DVCon.
- Poulos and Veneris, Clustering-based Failure Triage for RTL Regression Debugging, ITC 2014.
- Luu et al., VCDiag: Classifying Erroneous Waveforms for Failure Triage Acceleration, 2025.
- Automating Regression Triage in Design Verification Using AI-Based Random Forest Models, DVCon.
Comments (0)
Leave a Comment