AI Coverage Closure: What Actually Closes Bins

Every claim in the vendor decks is true, and most of them are not about closing coverage. That is the puzzle this post untangles. AI tools for coverage closure are real, deployed, and in a few verified cases spectacular — one production SoC interconnect went from a 79.91% plateau to 100% functional coverage at 18x less compute. But the marquee numbers you have seen — 16x, 10x, 5x — are almost all compression results: the same coverage, reached cheaper. Closing new coverage is a different problem, the published wins at it are rarer and more conditional, and the conditions are the interesting part. This post maps the manual closure loop you already run, sorts the AI offerings into three tiers by what they actually touch, walks through the verified production results, and then spends equal time on what none of them automate — because the honest boundary of these tools is exactly where your judgment still lives.

The Last-Mile Problem

Start with the loop as every DV team actually runs it, stated with unusual candor by an NVIDIA team at DVCon: identify covergroups, code the bins, run regressions, generate reports and analyze, "modify constraints manually and rerun the regressions," and — step six — "repeat steps 4 and 5 until we hit 100%." Behind that "repeat" sits the whole machinery: merge the coverage databases (urg, imc, vcover merge), rank the tests, triage the holes, categorize each one — needs a new test, needs a constraint tweak, genuinely unreachable, or waivable — write exclusions, get them reviewed, rerun. The loop is not broken. It just has a cost curve that turns hostile precisely when you need it most.

The best production dataset ever published on that curve comes from a Samsung + Cadence paper on a mobile application processor's SoC interconnect — hundreds of masters, near a thousand slaves, hundreds of thousands of functional coverage elements (DVCon US 2025). One regression pass: 274 runs, 1,744 CPU-hours, 26.74% coverage. Twenty iterations: 5,480 runs, 33,751 CPU-hours, 94.84%. The next four iterations — 1,096 more runs, 7,408 more CPU-hours — bought 1.43 points. Do the division and the efficiency collapse is about 15x: the first 94.84% cost roughly 356 CPU-hours per coverage point; the tail cost about 5,180. After 41,158 CPU-hours, 4,478 bins were still open.

A software engineer will recognize the shape instantly: it is a feedback loop with exponentially decaying reward. Constrained-random stimulus re-hits the easy bins the way CI re-runs already-passing tests — every regression pays full price to reconfirm what the last one proved, and the probability of landing on an unhit bin shrinks as the unhit set does. Infineon quantified the redundancy on a production radar DSP in their Aurix line: closing coverage took on the order of two million tests, a thousand machines and licenses running nearly continuously for six months — and an industry-standard ranking algorithm afterward showed that about 3,000 of those tests would have sufficed to hold 100% coverage (DVCon US 2021). Over 550 tests simulated for every one that added coverage.

This is the pressure the AI tools are selling into, and the pressure is real. The 2024 Wilson Research Group study — still the latest as of this writing — has first-silicon success at 14%, the lowest in two decades of tracking; 75% of IC/ASIC projects behind schedule; design engineers spending 49% of their time on verification. Nobody needs convincing that the loop is expensive. What needs examining is which part of it each tool actually touches.

Your Regression Is 30–60x Bigger Than Its Coverage

Before any AI enters the picture, understand what plain test ranking already proves, because it is the baseline every ML claim should be measured against — and in the one independent study that did measure, the baseline nearly tied.

Ranking is the greedy algorithm your coverage tools have shipped for years: given a merged, test-associated coverage database, select the minimal set of test-seed pairs that reproduces the merged coverage. Infineon ran it across three real projects with Cadence (DVCon) and the numbers are startling: a microprocessor IP regression of 260 runs compressed to 8 runs at 100% coverage regain — 32.5x. A mixed-signal SoC's 5,124 runs compressed to 1,204, again at full regain. Your regression is 30–60x larger than its minimal coverage-proving subset, and the tooling to prove it is already in your license.

Two caveats turn this from a party trick into engineering judgment. First, ranking hits exactly the bins the original regression hit — never one more. It is compression by construction, and a compressed regression cannot close a hole. Second — and this is the finding teams skip past — the same study showed the compressed regressions changed the failure profile: one stage went from 3 failing runs in the original to 23 in the optimized set. The redundant tests were not worthless; they were soak. The coverage-proving subset and the bug-finding regression are different objects, and a tool that optimizes the first while you silently assume it preserved the second is how a "more efficient" flow ships a bug.

When the same study benchmarked Cadence's Xcelium ML against this non-ML ranking baseline, the result was deflationary and useful: "Both Xcelium ML and Ranking methods gave comparable compression & speedup factors around 3 consistently" — with ranking sometimes compressing more. The ML tool's one structural advantage: its regenerated regressions occasionally exercised genuinely new scenarios, regaining more than 100% of the original coverage (101–108% in several configurations) — something ranking cannot do by construction. Hold that thought; it is the entire difference between the next two sections.

One more piece of the baseline: hole triage has five buckets, not four. Needs a new test; needs a constraint or seed change; genuinely unreachable (route it to formal, not to a human); waivable with review; and — the bucket most flows don't have — the coverage model itself is wrong. A Samsung Memory team measured that last bucket on a production cache-managing IP: 14.1% of their coverage holes were bins that were never defined at all — the initial hand-written model covered only 53.5% of the true bin space (DVCon). Keep that number in mind when a tool promises to close your holes. Some of your holes are not holes.

Three Tiers of "AI Coverage" — Read the Fine Print

Every commercial "AI coverage" offering does one of three things, and the tier determines what the tool can possibly deliver. Sorting the market this way is the single most useful filter you can apply to a vendor deck.

Tier A — selects or ranks existing tests. Same stimulus pool, fewer cycles. AMD's SNUG result with Synopsys VSO.ai — 1.5–16x fewer tests to the same coverage across four designs — lives here, as do Renesas's 2.2x/3.6x Xcelium ML compressions, VSO.ai's regression-ROI ordering, and the change-based smoke-suite selection Intel presented in 2026. A Tier A tool, by construction, cannot close a coverage hole. It can only make the coverage you already reach cheaper — which is genuinely valuable, and is where nearly every marquee number comes from.

Tier B — steers the constrained-random distribution. The tool reaches into the randomization kernel or constraint solver and re-weights what stimulus gets generated, from the same testbench and the same constraints. Xcelium ML's accelerated-closure mode, Cadence's Verisium SimAI, and VSO.ai's in-simulator coverage-directed solving live here. Tier B can hit bins that plain random rarely reaches — this is where the real closure results live — but it cannot reach anything your constraints exclude. It explores your legal space more cleverly; it does not enlarge it.

Tier C — authors new verification artifacts. Tests, sequences, properties that did not exist before. As of mid-2026 no dedicated coverage-closure product from the big three is in this tier. What is here: the research wave (agentic property generation, LLM testbench synthesis) and the vendors' new agentic layers — Cadence's ChipStack, Synopsys's AgentEngineer, Siemens's Questa One Agentic Toolkit, all announced within a single month in early 2026, all early-access, all pitched at bring-up and productivity rather than last-mile closure.

The reader's rule that falls out: for every number in a vendor deck, ask did coverage go up, or did the same coverage get cheaper? Renesas's split is also worth carrying with you — 3.6x compression on a derivative design versus 2.2x on the original, because the ML had regression history to learn from. These tools eat your data, and a team with deep regression archives has a moat a fresh project does not. As one DAC 2026 wrap-up put it, the EDA companies build the tools, but the training data belongs to the companies building chips.

What Actually Closed New Coverage

Now the wins that survive fact-checking — each one primary-sourced, each with its condition attached.

The Samsung + Cadence interconnect result is the strongest closure number in public. On the second covergroup category — the one where the traditional flow plateaued at 79.91% after 41,158 CPU-hours — the SimAI-guided flow reached 100% in 2,261 CPU-hours: 18.2x less compute and 20.09 points more coverage, from the same testbench. On the first category, where the baseline had already clawed to 96.27%, the gain was 4.92x. Read those two numbers together: the advantage was largest where the tail was worst. The tail is where these methods earn their keep, not where they fail. ("Up to 18x" is doing real work in the abstract, though — the two categories are the whole spread.)

A second Samsung team reported the same shape on a production camera-interface IP: 134 cross covergroups, 14,616 bins; the ML-driven flow reached 100% in about 75,000 tests while the baseline regression "fails to exceed 70% coverage rate, even after more than 100,000 runs."

Infineon's novelty-driven selection is the best-documented Tier A result: an autoencoder ranks candidate tests by reconstruction error — novelty, in effect anomaly detection pointed at your own stimulus — and simulation order follows novelty. 60% fewer tests to reach 99.5% coverage, still 40% fewer at 99.95%, projecting the six-month closure campaign to under three months. Two honesty flags the paper itself carries: the time saving is a projection from an offline replay of 85,470 already-generated tests, not a measured deployment; and the savings decay from 60% to 40% exactly in the final half-percent, where the hard bins live.

NVIDIA's VSO.ai deployment — designs exceeding 100 million coverage targets — reported 33% more functional coverage in the same number of runs, alongside a 5x regression-suite reduction. Their published flow combines test grading, formal unreachability analysis, and VSO.ai; the often-quoted 17%-more-coverage-at-3.5x-compression figure belongs to that combined flow, not to the AI tool alone. Their lesson learned is also on the record: the tool was first tried late in a project, and the vendor now recommends deploying at early milestones, while stimulus is immature — late-stage deployment underperforms.

On the research side, one result is worth more than all the benchmark tables: LLM4DV, the Cambridge/Imperial/lowRISC benchmark for LLM-driven stimulus generation (FCCM 2025). In its 2023 version, GPT-3.5 managed 5.61% coverage on the hardest DUT, an Ibex CPU. In the current version, on identical scaffolding, Claude 3.5 Sonnet reaches 100% on that same CPU — against a constrained-random baseline of 15.31%. Across the benchmark the per-model spread runs from 7.93% to 98.84% on the same design with the same framework. The framework was never the binding constraint. The model was. Whatever you concluded about LLM stimulus generation from a 2023-era evaluation, the conclusion has a shelf life measured in model releases.

The Catch: A Human Still Writes the Targets

Here is the pattern connecting every win above, and it is the thesis of this post: the automation is in reaching specified targets efficiently, not in discovering what to target. The coverage model and the scenario list remain human artifacts, and every failure mode in this section is a way of forgetting that.

The Samsung + Cadence paper says it outright: the DV team supplies the target scenario specification, and "if the target specification is incomplete… the proposed approach may not hit the bins even though all the test scenarios of the provided specification are satisfied." Deriving the specification automatically is listed as future work. The 18x result is a solver being steered brilliantly toward targets a human enumerated.

Nokia and MathWorks hit the boundary from the other side (DVCon Europe 2023). Their autoencoder test selection cut tests-to-closure by up to 43% — but the coverage goal was capped at 67%, "the maximum possible for the current test randomization constraint configuration." No selection strategy, however intelligent, could touch the remaining third, because the constraints excluded it. Selection cannot fix a constraint problem. (The paper's abstract says "2x speedup"; its own results section reports that wall-clock regression time was not improved — doubled, in the worst case — because each iteration restarted the simulation environment. Cite this paper carefully.)

The Samsung missing-bins result gives the model-side version: 14.1% of holes were bins nobody wrote. An AI aimed at your holes is aimed at the defined bin set. The gap between the bins you wrote and the bins you should have written is invisible to every tool in this post — as the Verilab crew put it in the best paper ever written on coverage quality, "if something is missing from the model, it does not appear as a coverage hole, it's simply invisible." Their companion warning belongs on a wall: functional coverage modeling "is essentially a software task. As such, models will likely have bugs… Given the trust we put in functional coverage results for tapeout decisions, this is an oddly overlooked requirement." Mark Litterick's Lies, Damned Lies, and Coverage names the failure taxonomy — deception, omission, fabrication — and the reason coverage bugs survive: "if you make a mistake in stimulus or checks, there is a good chance you will kill the regression; if you mess up coverage there is no comeback." Goodhart's law — when a measure becomes a target it ceases to be a good measure — comes from economics via the software-testing literature, but DV built its own sharper version first.

Now put an AI in that loop and watch the metric detach from the goal. Infineon's agentic formal-coverage work (arXiv:2603.03147) is admirably honest about what happened: LLM agents read Jasper coverage reports, characterized the uncovered RTL, and generated new SVA properties, lifting formal coverage by roughly 10–20% across five designs. And: "in some cases… the proven rate for generated properties decreased after the coverage agents' workflow, even though overall coverage increased." The coverage number went up while the fraction of properties that could actually be proven went down — and the published results ran with no human review in the loop. That is the metric improving while assurance does not, measured and self-reported.

The sharpest 2026 datapoint on where AI stimulus generation actually stops comes from a hole-by-hole taxonomy of everything an agentic flow failed to close across 19 designs (arXiv:2604.15657). Under 7% of the residual holes were genuinely unhittable — tied-off integration logic, defensive dead code, infeasible boundaries. About 92% were reasoning frontiers: multi-module pipeline warm-up sequences (49.9% of the frontier bucket) and protocol sequencing (40.2%) — holes that require building a Wishbone burst model or an MDIO responder to reach. The authors' key finding: "the agent correctly diagnoses these problems, [but] it fails to implement the solutions." The agent can read a coverage report and explain the hole like a staff engineer; it cannot yet write the protocol machinery to hit it. Diagnosis has been automated. Generation, at protocol depth, has not.

So the reformulated claim that survives all the evidence: AI earns its keep on tail bins that are reachable and correctly specified. The unreachable and the unenumerated stay yours.

Formal UNR: The Workhorse With Its Own Wall

The least glamorous automation in this story predates the AI wave and out-delivers most of it. Unreachability analysis takes your partial coverage database plus the RTL, formally proves which uncovered targets cannot be hit under any stimulus, and emits an exclusion file back into your coverage flow — Jasper's UNR app, VC Formal's FCA invoked natively from the VCS shell, Questa CoverCheck. A Questa team's DVCon tutorial reported a PCIe-bench UNR run of three hours that saved an estimated three weeks of manual analysis; Synopsys's blog cites Cisco seeing a 9% coverage improvement from pruning noise. One design in that same tutorial had over 3,000 unreachable coverage elements — at even 15 minutes of human triage each, that is 4.5 person-months of analysis a formal engine did before lunch.

Two things keep this section honest. First, nobody has a consistent answer to "how much does UNR save": Synopsys's own materials claim 40–80% verification-effort savings in one publication and 8–80% in another. The technique is real; the aggregate number is marketing. Second, a Qualcomm engineer's account from a VC Formal SIG punctures the assumption that UNR is a solved deployment: on their larger configuration — over 200 million coverage goals — the analysis produces claims at a scale where "millions… may need manual review by design experts, an impractical expectation only manageable through engineer-defined blanket exceptions which undermine the integrity of the analysis." And the exceptions don't port across configurations or even successive RTL drops. UNR "still is not as mainstream as you might imagine."

One reframe worth stealing from Synopsys's FCA documentation: an unreachable target you expected to be reachable is not a waiver candidate — it is a design bug wearing a coverage costume. And note that UNR-versus-ML is a false rivalry: NVIDIA's published flow runs test grading, UNR, and VSO.ai together. The formal engine prunes the impossible; the ML steers toward the merely improbable.

Exclusions Are Code

Every closure flow — manual, formal, or AI — terminates in the same artifact: an exclusion list that redefines what 100% means. Treat that artifact with the same rigor as RTL, because it carries the same tapeout risk. The best public, enforced example is OpenTitan's DV methodology:

  • Every exclusion carries a standardized annotation prefix — UNR, NON_RTL, UNSUPPORTED, EXTERNAL, LOW_RISK — making the exclusion base greppable and auditable. LOW_RISK is the honest bucket: a named, reviewed "we chose not to chase this," instead of a silent one.
  • Designers sign off on exclusions in PR review — the person who wrote the RTL certifies the code is genuinely uncoverable, not the person whose schedule benefits from the waiver.
  • "If any RTL changes happen to the design after the coverage exclusion file has been created, it needs to be redone and re-reviewed." Exclusion work starts only after design freeze, precisely because of this rule. The Questa tutorial gave the failure mode a name worth adopting: waiver rot — manually generated waivers "have to be maintained as the code changes."
  • Coverage is one line item among many: the V3 signoff gate also requires all assertions proven, no unreachable properties, and a nightly regression 100% passing with a week of soak. Closure is a portfolio, not a number.

The AI angle lands directly here. At least one new-entrant "coverage agent" product's headline capability is recommending exclusions with supporting evidence. Read that plainly: it closes coverage by removing targets. That may be exactly right — a good UNR-plus-triage assistant is valuable — but an AI-recommended exclusion must enter the same gate as a human one: annotated, designer-signed, invalidated on RTL change. An agent that can edit your exclusion file has write access to the definition of done.

Build Your Own Loop: The API Surfaces That Matter

Suppose you want the loop the papers describe — a model reading coverage, choosing what to run next — without waiting for a product. What can you actually build against, today? The answer has a clean structure, and it starts with zero APIs at all.

The loop you can write this afternoon lives entirely inside IEEE 1800. Every covergroup, coverpoint, and cross has a get_coverage() method, and pre_randomize() is the standard-blessed hook that runs before every randomize() call. Put them together and the testbench biases its own stimulus toward whatever is least covered:

function real calc_weight(opcode_t op);
  real cov;
  case (op)
    nop_op:  cov = covunit.cg.op_nop.get_coverage();
    load_op: cov = covunit.cg.op_load.get_coverage();
  endcase
  return (100 - cov) * 0.5;   // colder coverpoint => heavier weight
endfunction

function void pre_randomize();
  weight_nop  += calc_weight(nop_op);
  weight_load += calc_weight(load_op);
endfunction

No vendor dependency, works on all three simulators, and it is a real coverage-feedback controller — a proportional controller, in control-theory terms. It also teaches you the standard's load-bearing limitation by running into it: the LRM gives you coverage percentages, never bin identities. §19.9 defines exactly three covergroup system tasks ($set_coverage_db_name, $load_coverage_db, $get_coverage) plus the get_coverage methods; there is no standard way to ask which bins are empty from inside a running simulation. Everything more ambitious than proportional weighting needs the coverage database — which means it happens between runs, not during them.

The coverage database is the real interface, and one vendor documents it. Siemens publishes the full Questa UCDB C API — a 200+ page reference with worked examples shipped in the install — and the UCDB test record turns out to already be an ML training record: it stores the test name, the seed (verbatim from -sv_seed), a command field the docs describe as capturing "knob settings for parameterizable tests," CPU and simulation time, and pass/fail status. Stimulus knobs, seed, cost, outcome — the (action, cost, reward) tuple, persisted by the tool you already run, surviving merges. The Tcl layer above it is equally direct: vcover merge -testassociated (nothing per-test works without it), then coverage analyze -select cover -eq 0 — hole extraction as a query — then coverage ranktest, which emits ranktest.contrib and ranktest.noncontrib: machine-parseable lists of contributing and redundant tests, which is to say, labeled training data for a compression model, generated by a shipping tool. There is even a sanctioned place for an agent to stamp its own metadata into the database: coverage attribute -trendable.

The other two vendors gate their equivalents — Synopsys's coverage C API manual is marked Confidential; Cadence's IMC reference and vManager REST API live behind support portals. The practical trick: open-source consumers are the real documentation. OpenTitan's production fpv.tcl is a better JasperGold coverage reference than anything public from Cadence (check_cov -init before design load, -measure with a time limit, -report to parseable output — and waivers are just a Tcl file, which means an agent's exclusions are just a file it writes, subject to the previous section's gate). The Jenkins vManager plugin documents the vAPI REST surface by using it. And imc -execcmd "help report" makes the tool document itself.

On machine-readable output, the ground truth is humbling. The VCS manual documents exactly two urg output formats — text and HTML; zero occurrences of XML, JSON, or CSV. Cadence's flow, per OpenTitan's production scripts, emits text summaries and HTML (plus a native rank command). Across the industry, the practical egress contract for coverage data is parsing text reports — which retroactively explains the most striking pattern in the published work: neither industrial AI-closure paper used a vendor API. Nokia shelled out to ExecMan and parsed results; Infineon's agents shell out to Jasper and parse reports. The standard that was supposed to fix this — Accellera UCIS — was ratified in 2012 and never revised; its working group is inactive, no vendor publicly documents a conformant UCIS shared library, and the most credible open implementation (pyucis) had to write its own. The working loops route around the standard: Infineon's ISCAS 2025 flow skipped database extraction entirely and had the PyVSC coverage callback in a PyUVM monitor append stimulus values and bin hit/miss flags to a CSV during the run. Their stated reason is the thesis of the open-source path: "PyUVM testbenches offer a significant advantage in data collection compared to SystemVerilog-UVM testbenches." Python's edge is not nicer constraints. It is data egress — the testbench already lives in the language the ML lives in.

For tools with no API at all, two bridge patterns cover everything. Inside the simulation: DPI-C plus a TCP socket — both published RL closure loops (a DQN closing a compression encoder's CAM bins, and Infineon's Gym-based agents) independently built the same bridge, SystemVerilog importing a DPI function whose C side talks to a Python agent over a socket. One discipline point if you try it: the foreign call blocks the simulator, so make decisions in pre_randomize() or between transactions — a model consulted inside a sampling path perturbs timing and breaks seed-stability. Outside the simulation: a Tcl socket listener running inside any Tcl-shelled tool (vsim, ucli, Jasper, IMC) with a thin MCP server outside — a pattern already demonstrated by a third-party Xcelium debug server, and generalizable to every tool in your flow. No vendor cooperation required.

And the economics lesson that decides whether any of it pays: Nokia's loop reduced simulated tests by 43% and still failed to improve wall-clock time — worst case, regression time doubled — because every iteration tore down and restarted the simulation environment. The integration point that determines whether an AI loop is economical is not the coverage API. It is whether you can keep a warm simulator between decisions. Design for that first.

One boundary to scope your ambitions honestly: the commercial Tier B tools plug into the constraint solver and randomization kernel — a layer of the stack with no public API on any simulator. You cannot build VSO.ai or SimAI from outside. Everything upstream of the solver (test selection, knob choice, seed allocation) and downstream of it (hole extraction, triage, ranking, property generation, exclusion drafting) is buildable today with what this section named.

Piloting Without Getting Burned

If the post has a single operating principle, it is Mike Bartley's: separate productivity gains from assurance gains. A tool that halves your regression bill has improved productivity; whether your verification got better is a different question with different evidence, and the Infineon proven-rate result shows how easily a coverage metric can rise while assurance falls. His pilot guardrails compose into a checklist that fits on an index card: bound the pilot's scope; define success in engineering terms before deployment; version your datasets with traceability to DUT revision — "if the regression environment cannot reliably map failures to DUT revision, scenario, configuration, and known bug state, the model will learn noise"; route generated artifacts (tests, properties, exclusions) through the same review path as human-authored ones; and keep sign-off authority human.

Add the deployment lessons the case studies paid for: deploy at early milestones, not late (NVIDIA's late-project trial underperformed; the vendor now says so). Expect results proportional to your regression history — the derivative-design effect is real, and a team with years of archives will see numbers a fresh project will not. And run the trial on your design with a defined baseline, because the independent-evaluation landscape is thin enough to be its own finding: for the most heavily marketed tool in this space, every public number still routes through the vendor, and the one independent multi-project study found the ML tool roughly tied with plain seed ranking on its headline metric. That is not a reason to skip these tools. It is a reason to measure them the way you would measure anything else you were about to trust with tapeout: on your data, against your baseline, with the metric and the goal kept honestly apart.

The last mile of coverage closure has always been where verification stops being mechanical and starts being judgment — deciding what the model should contain, what the constraints should allow, what the design can never do, and what you are willing to sign. The verified wins in this post are real, and none of them moved that boundary. They cleared the brush on the road to it. Walk the last stretch yourself, and know exactly where it starts.


This post is part of the AI for DV series. Claims above trace to the primary sources linked inline; the research corpus and verification notes live with the series. Related: the Practitioner Playbook and the AI Reading List.

Author
Milan Kubavat
Sharing knowledge about silicon verification, hardware design, and engineering insights.

Comments (0)

Leave a Comment