Better Prompts for DV: A Researcher Guide to Prompt Engineering in Hardware Verification

Most prompting research is generic. Most DV blog content about AI is hype. This post sits in the small intersection: a researcher guide to prompt engineering techniques mapped explicitly to Design Verification problems, with research callouts naming the papers, numbered playbook patterns as the practitioner takeaway, and DV-application blocks grounding each technique in your UVM testbench or RTL workflow.

The differentiator is Section 10 — the systematic correlation between SWE prompting strategies (where the empirical research has been done) and the DV adaptations they imply. Section 12 closes with six ready-to-paste prompt templates, each annotated with the research pattern it borrows from.

A note on what this post is not. It is not a survey of LLM model capabilities (those change quarterly). It is not a tutorial on Anthropic vs OpenAI vs your local model (the workflow is provider-agnostic). It is not advocacy for AI replacing DV engineers (the research consistently shows the human-in-the-loop is what wins). It is the durable methodology layer beneath whichever model and provider you happen to use today.

1. Foundations: Zero-shot, Few-shot, Role

Before the fancy techniques, the basics that ground everything else. Each of the three primitives below has a clean role in DV; mixing them up is a leading source of frustration.

Research arxiv 2406.06608 The Prompt Report — PRISMA-grounded systematic survey of 58 prompting techniques. arxiv 2402.07927 Systematic Survey of Prompt Engineering. The 2025 MDPI review of 42 peer-reviewed SE studies clusters prompting into four research patterns: manual crafting, RAG, chain-of-thought, automated tuning. The Liu et al. 2026 taxonomy formalizes 58 distinct techniques across modalities.
  1. Zero-shot. A single instruction, no examples. Right for: tasks the model has clearly seen many times (generate a docstring, summarize a paragraph, classify by sentiment). Wrong for: anything that involves your team's idioms, your codebase's style, or a non-standard output schema.
  2. Few-shot in-context learning. 2-5 worked examples before the actual ask. The model picks up the shape of the desired output. Critically: examples should come from your own codebase, not from a generic library — the model already knows generic libraries.
  3. Role-prompting. "You are a senior DV engineer reviewing a 10-year-old testbench." The role routes the model into a more conservative, idiom-aware mode than the default helpful-assistant frame. The 2025 SE literature notes role prompts can be removed without much loss once you have good few-shot examples; until then, they are doing real work.
Apply to DV
  • Zero-shot for repetitive boilerplate: covergroup skeletons, factory registration, default field automation macros.
  • Few-shot for anything style-sensitive: your team's sequence library idioms, your scoreboard pattern, your monitor structure.
  • Role for tasks where the default helpful-assistant tone would be wrong: code review (assistant is too lenient), architectural critique (assistant defers), and triage (assistant hedges).

2. Chain-of-Thought for DV: The Caveat Section

CoT is the most famous prompting technique. It is also the one most likely to disappoint you on RTL tasks if you apply it naively. The fix is to scaffold the reasoning with HW-meaningful intermediate steps, not generic "let's think step by step."

Research Wei et al. 2022 (arxiv 2201.11903) — CoT activates reasoning in large models; without it even 540B-parameter models behave like much smaller ones. The 2026 RTL survey (Preprints 202509.1681) explicitly states that naive chain-of-thought has been largely ineffective in automating IC design workflows because the step granularity and reasoning direction do not align with expert RTL knowledge. arxiv 2504.06939 (FeedbackEval) shows that for code repair, removing structured-reasoning cues causes severe degradation.
  1. HW-aware CoT scaffolding. Replace "think step by step" with "walk through this cycle by cycle." The model needs the right intermediate representation; for hardware that is signal traces, FSM transitions, timing windows, and clock-domain crossings — not generic prose.
  2. Trace-conditioned CoT. If you have a waveform or structured log, paste the relevant signal slice into the prompt before asking for analysis. arxiv 2505.04441 shows trace-conditioned prompts consistently beat trace-free prompts for SWE code repair; the DV analog is direct.
  3. Bounded CoT. Ask for reasoning in N steps, not free-form. "Identify the failing cycle. Identify the violating signal. Identify the upstream cause. Propose the fix." Four bounded steps beat an open-ended "think about this."
  4. Avoid CoT for structural generation. RTL module generation rarely benefits from CoT — the model needs to emit a structured artifact, not reason about it. Use decomposition (Section 5) instead.
Apply to DV
  • HW-aware CoT for timing-related debug: assertion failures involving multi-cycle properties, arbitration races, CDC investigations.
  • Trace-conditioned CoT for any failure where the structured log captures the relevant signal evolution — combine with the JSON ring-buffer dump pattern.
  • Skip CoT entirely when asking for boilerplate or scaffolding; you want the artifact, not the model's commentary on producing it.

3. Self-Consistency and Self-Verification

A complex problem usually admits multiple correct reasoning paths. Sample many paths, vote on the consistent answer. The technique is dramatic on hard problems and almost free when the model supports temperature sampling.

Research Wang et al. 2022 (arxiv 2203.11171) introduces self-consistency: sample diverse reasoning paths, take the most consistent answer. arxiv 2510.01069 Typed Chain-of-Thought / Certified Self-Consistency (CSC, 2026) aggregates only over experiments satisfying typing constraints — 69.8% accuracy on GSM8K versus 19.6% baseline. arxiv 2603.08999 Confidence-Aware Self-Consistency maintains comparable accuracy with up to 80% fewer tokens. Self-Verification (Weng et al. 2022, arxiv 2212.09561): generate forward, verify backward.
  1. Vote on hypothesis-rank. When triaging a critical bug, run the hypothesis-rank prompt three times with temperature > 0. Take the hypothesis that appears highest in all three runs. Cheaper than it sounds; catches obviously-wrong single-shot answers.
  2. Verify forward, check backward. Generated a fix proposal? Ask the model to derive the test that would fail without the fix. If the model cannot, the proposal is suspect.
  3. Certified self-consistency for typed outputs. When the answer must conform to a schema (a coverage plan, a JSON fix proposal), aggregate only over candidates that pass type/schema validation. The 2026 CSC paper is the formal version of what good DV teams already do informally.
  4. Confidence-aware sampling. When the first two samples agree strongly, stop. When they disagree, sample more. The 2026 CASC paper formalizes the trade-off; the practical takeaway is to avoid blindly running N samples when 2-3 suffice.
Apply to DV
  • Self-consistency on RCA for silicon-escape bugs — the stakes justify the extra samples.
  • Self-verification on every proposed RTL fix before the human reads it: ask the model to derive a failing test for the unpatched version.
  • Certified self-consistency for coverage-plan generation: aggregate only over plans that parse against your covergroup schema.

4. Tree-of-Thoughts for Hard RCA

When chain-of-thought gets stuck on a wrong path, Tree-of-Thoughts explores multiple branches in parallel and prunes. The headline result is dramatic; the DV use case is hard root-cause analysis where 3+ hypotheses are equally plausible.

Research Yao et al. 2023 (arxiv 2305.10601) introduces Tree-of-Thoughts. The headline: on Game-of-24, GPT-4 with standard CoT solved 4% of problems; with ToT, 74%. arxiv 2401.14295 (Besta et al., Demystifying Chains, Trees, and Graphs of Thoughts) extends to graph-shaped reasoning. The pattern works because branching defers the commitment to a single reasoning trajectory.
  1. Branch the hypothesis space. "Generate three independent root-cause hypotheses, each starting from a different evidence anchor in the bundle." Force the model to start from different observations, not extend one chain of reasoning.
  2. Score and prune. After branching, ask the model to score each branch by "evidence weight" and propose which to prune. The model is often better at evaluating its own branches than picking one upfront.
  3. Iterate on the survivor. Take the highest-scored branch and apply CoT or self-debug only to that subtree. Saves the cost of exploring all branches deeply.
  4. Use when bug is "ambiguous." If a senior engineer cannot quickly name the most likely root cause from inspection, that ambiguity is the trigger for ToT. For obvious bugs, vanilla hypothesis-rank is cheaper.
Apply to DV
  • ToT on multi-IP SoC bugs where the failure could plausibly involve cache, fabric, or peripheral — let the model branch by subsystem.
  • ToT on protocol-error bugs where the violation could be the master, the slave, or the interconnect — branch by node.
  • Skip ToT for clear-cut bugs — the cost is real and the marginal benefit is zero when the answer is obvious.

5. Decomposition: Least-to-Most and Decomposed Prompting

Complex problems become a sequence of simpler ones. The technique transferred wholesale from theorem-proving to LLM prompting and remains one of the highest-leverage moves in your toolkit for any task larger than a single function.

Research Zhou et al. 2022 (arxiv 2205.10625) Least-to-Most Prompting — 16% to 99% on SCAN compositional generalization. arxiv 2210.02406 Khot et al. Decomposed Prompting — modular subproblems with composable solvers. The 2026 RTL research (PALM, arxiv 2506.09002) extends decomposition to program-analysis-derived path constraints; the DV analog uses coverage-bin-derived path constraints.
  1. Stage 1 — decompose. Ask the model to list the subproblems before solving any. "To verify this IP, list the 6-8 distinct sub-tasks in order of dependency." The decomposition itself is high-value output.
  2. Stage 2 — solve in order. Each subproblem solution becomes context for the next. The model maintains coherence because each step is bounded.
  3. Dependency-aware decomposition. When subtasks have non-linear dependencies, draw the DAG. The model can produce the DAG and then walk it in topological order.
  4. Path-constraint decomposition. For test generation: derive path constraints from coverage bins (analog of program-analysis-derived branching conditions in PALM), use each constraint as a sub-prompt.
Apply to DV
  • Verification plan: spec → coverage plan → sequence plan → scoreboard plan → checker plan → test list. Each step a separate prompt, each consuming the prior.
  • UVM agent scaffold: interface → transaction class → driver → monitor → sequencer → agent → example sequence. Each generated separately, each consistent with the prior.
  • SoC verification: top-level constraints → per-IP test plans → integration scenarios → stress tests. Top-down decomposition mirrors how senior engineers actually plan.

6. Few-Shot Done Right: Example Selection

Few-shot is the most common technique and the most commonly done badly. Random examples are sub-optimal. Text-similarity examples are sub-optimal. The 2024-2026 research has converged on better methods.

Research arxiv 2310.09748 LAIL (LLM-Aware ICL for Code Generation) — LLM labels candidate examples as positive (helpful) or negative (trivial); a model-aware retriever learns the preference. arxiv 2305.14210 Skill-Based Few-Shot Selection — pick examples that share the underlying skill needed, not just textual surface. arxiv 2412.02906 empirically: few-shot helps code synthesis but example quality dominates count.
  1. Examples from your codebase, not from the public web. The model already knows the public web. Your codebase carries your idioms, your naming, your error-handling conventions; the model picks these up from few-shot context cheaply.
  2. Structurally-similar examples beat textually-similar ones. Two AXI sequences that look different on the surface but share the same burst structure are better few-shot fodder than two cosmetically similar sequences with different structures.
  3. Negative examples are valuable. "Here is a similar task done wrong, here is the correct version." The contrast teaches the model what to avoid; bare positive examples cannot.
  4. 2-5 examples, no more. Beyond five, the marginal benefit per token drops sharply and the long-context degradation effect (arxiv 2509.21361) starts to bite.
Apply to DV
  • Maintain a curated examples directory in your TB repo: examples/sequences/, examples/scoreboards/, examples/checkers/. Each example is a known-good reference you few-shot from.
  • Index examples by protocol family and skill (AXI burst, AXI single, PCIe TLP, USB packet) so retrieval is structural.
  • Keep a small set of "canonical wrong" examples for negative few-shot — the seq that deadlocked, the scoreboard that missed an off-by-one. Contrast teaches.

7. Agentic Prompting: ReAct and Reflexion

When the LLM needs to iterate — observe, decide, act, observe again — you are in agentic territory. ReAct and Reflexion are the foundational patterns; the 2025-2026 HW research extends them with verifier-guided refinement.

Research Yao et al. 2022 (arxiv 2210.03629) ReAct — interleave Reasoning steps with Action calls (tool invocations) and Observations. Shinn et al. 2023 (arxiv 2303.11366) Reflexion — the agent self-critiques after each task and improves next time. arxiv 2509.06239 Proof2Silicon — verifier-guided prompt refinement via reinforcement learning, applied to verified hardware code generation.
  1. ReAct loop. Reason → Act (call a tool: query log, run smoke, fetch RTL excerpt) → Observe (tool output) → Reason. Three to five iterations beats one zero-shot answer almost universally for non-trivial problems.
  2. Bounded tool catalog. Per the 2025 Anthropic context-engineering essay: 4-6 self-contained tools beat one mega-tool. Tools must be self-contained, error-robust, and unambiguous. If a human cannot decide which tool to call, the agent cannot either.
  3. Reflexion after task completion. After the loop converges (or fails), ask the agent to write a paragraph about what worked, what did not, what to try differently next time. Cache the reflection; use it in the next session's system prompt.
  4. Verifier-guided refinement. Per Proof2Silicon: when a verifier (smoke test, formal property, lint rule) gives concrete feedback, route it back into the prompt as a hard signal. This is the HW-side analog of test-driven debug from DePro (arxiv 2603.19399): up to 64% fewer attempts and 7.6 minutes saved per problem in the SWE setting.
  5. Checkpoints and human approval. Per the 2025-2026 agentic-AI consensus: free-form agent loops are less reliable in production than graph-based orchestration with explicit state transitions, debuggability, and human-approval gates. Plan for the human checkpoint; do not assume autopilot.
Apply to DV
  • ReAct for triage: query_log, get_rtl_excerpt, run_smoke, query_fingerprint_db as the tool catalog. Anything else is scope creep.
  • Reflexion for postmortem: have the agent draft the "what we learned" section from its own trace. Edit for tone before publishing.
  • Verifier-guided refinement when you have a formal property: let the agent iterate on the property/fix until the formal tool stops complaining. The HW-side feedback signal is unusually strong.
  • Checkpoint on every fix proposal that touches RTL — never let the agent commit autonomously, regardless of how confident it sounds.

8. Constrained Decoding for Structured DV Output

When you need the LLM to produce a syntactically-valid artifact — JSON, Verilog, SVA, a coverage plan in your team's schema — you should be using constrained decoding rather than crossing your fingers. The infrastructure is cheap and the failure mode it eliminates is exactly the one DV engineers complain about most.

Research Constrained decoding modifies the LLM's sampling step via a logit processor: at each token position, valid tokens are computed from a grammar state and invalid tokens are masked. Libraries: llguidance (~50µs CPU per token; arbitrary context-free grammar), Outlines (compiles JSON schemas to O(1) lookup), Anthropic tool-call structured output, OpenAI JSON mode with response_format schemas. arxiv 2603.03305 Draft-Conditioned Constrained Decoding extends this to drafted generation with later refinement.
  1. JSON schema for structured outputs. Coverage plans, fix proposals, hypothesis-rank responses, test lists — all should land as schema-validated JSON, not as parseable-prose. The provider enforces the shape; you skip the defensive parser.
  2. Grammar-guided for HDL fragments. When asking for SVA or a small Verilog snippet, a grammar (or even a regex-shaped constraint) eliminates syntactic invalidity at the source. The model never emits the "almost valid" output that wastes a downstream tool invocation.
  3. Hybrid LLM + template. Per HAVEN (arxiv 2604.27643): the LLM produces a structured architectural plan in JSON; a rule-based generator emits the actual UVM. The split is what gets HAVEN to 100% compile success.
  4. Validate even with constrained decoding. Schema-validity is necessary, not sufficient. A schema-valid coverage plan can still be functionally wrong. Always pair constrained decoding with a downstream validation step.
Apply to DV
  • Define one JSON schema for each recurring DV output: coverage plan, fix proposal, hypothesis rank, test plan, refactor proposal. Reuse across teams.
  • For Verilog or SVA fragments, prefer hybrid generation: LLM produces a structured intermediate (signal list, property predicates, port map), template emits the syntactically-valid code.
  • Wire constrained-decoding errors into your alerting — if the model is failing constraint satisfaction repeatedly on a class of prompt, that is your signal the prompt itself needs work.

9. Meta-Prompting and Prompt Optimization

The prompts you write today are the next thing to refactor. Meta-prompting uses the LLM to rewrite your own prompts, A/B-tested against a small eval set. The 2026 industry consensus is that prompt engineering has matured from craft into versioned-and-tested engineering practice.

Research arxiv 2502.00728 Meta-Prompt Optimization for LLM-Based Sequential Decision Making. The 2026 industry literature (Comet, IntuitionLabs) reports self-refinement loops consistently improve prompt quality by 10-25%. arxiv 2503.02400 Promptware Engineering formalizes prompt engineering as a software engineering discipline: versioning, evaluation, regression testing for prompts.
  1. Build a small eval set. 15-25 representative DV prompts (debug, scaffold, review, refactor, cov-plan) with expected output shapes. This is your prompt regression suite.
  2. Self-refinement loop. Ask the LLM to critique and rewrite your prompt against the eval set. Compare A/B. Keep the winner. Repeat until the curve flattens.
  3. Version prompts like code. Prompts in a repo, semver tags, regression tests on every change. The day the underlying model updates, you re-run the suite to catch silent drift.
Apply to DV
  • Standardize the team's top 6-10 DV prompts in a shared repo. The hypothesis-rank prompt, the scaffolding prompt, the coverage-plan prompt — all versioned, all tested.
  • On every model upgrade or provider switch, re-run the eval suite. Catch the prompt that quietly stopped working before a junior engineer trusts a wrong answer.

10. SWE → DV: Methodological Correlations

This is the section that does not exist elsewhere on the internet. The SWE research community has done the empirical work on prompting for debug, test generation, code review, and refactoring. Each result transfers to DV with a clear adaptation. The table below names the SWE finding, the empirical evidence, the DV analog, and the specific adaptation needed.

Research Six recent SWE-prompting findings anchor this section. FeedbackEval (arxiv 2504.06939) on structured-reasoning code repair. DePro (arxiv 2603.19399) on iterative test-driven debug. PALM (arxiv 2506.09002) on program-analysis-derived path constraints. IntUT (paper covered in 2024-2025 SE literature) on intent-first unit-test generation hitting +94% branch coverage. Meta's semi-formal reasoning template hitting 93% code-review accuracy. Refactor subcategory explanation (arxiv 2411.02320) moving success from 15.6% to 86.7%. Static-analysis-augmented prompts (arxiv 2508.14419) cutting security violations from >40% to 13%.

The seven meta-patterns below each have an empirical SWE-side anchor and a direct DV adaptation. The DV adaptations are where the value is — they are not in the SWE papers because the SWE researchers were not thinking about hardware.

#SWE Meta-PatternSWE EvidenceDV Adaptation
1 Trace-Conditioned Prompting Execution traces in prompts consistently beat trace-free prompts (arxiv 2505.04441) Waveform-conditioned debug: paste FSDB signal slice formatted as "trace events"; Grove (arxiv 2511.x) confirms this for HW debug
2 Iteration-Feedback Loop DePro test-driven debug (64% fewer attempts); IntUT coverage-driven test gen (+94% branch coverage); FeedbackEval structured reasoning Smoke-driven self-debug loop; coverage-bin-driven test generation; formal-feedback loops (Proof2Silicon 2509.06239 is the HW-side prior art)
3 Intent-First Prompting IntUT: explicit test intentions improve branch coverage by 94%, line coverage by 49% Verification-intent-first: every UVM sequence prompt opens with "this is verifying scenario X / covering bin Y"; never "generate a test for this DUT"
4 Structured-Reasoning Templates Meta's semi-formal reasoning (premises → execution → conclusion) reached 93% code-review accuracy SVA-anchored review template: relevant assertions as premises; simulation/formal evidence as execution; coverage-bounded conclusion
5 Multi-Generation + Vote 5 refactorings per input → +28.8% pass rate; self-consistency over CoT samples Multi-fix-proposal + smoke gate: LLM proposes 3 fixes, smoke runs all 3, pick the one that passes. HW has a natural validator the SWE side often lacks — the simulator
6 Hybrid Tool+LLM Pipeline Static-analysis-augmented prompts (security >40% → 13%); PALM program-analysis path constraints Lint-augmented prompts; formal-augmented prompts; coverage-gap-augmented prompts. LASHED (arxiv 2504.21770) and Proof2Silicon are the HW-side prior art
7 Subcategory-Explanation Pattern Refactor success 15.6% → 86.7% from naming the refactoring type upfront (arxiv 2411.02320) UVM refactor categories upfront: extract base class, replace inheritance with composition, virtual-sequence extraction, agent split, factory-override simplification, config-object introduction. Name the category before asking for the change

Two threads tie the seven patterns together. First, every winning SWE prompting move has a feedback signal — tests, coverage, lint, formal — and the DV side has all of those signals available, often more cheaply than SWE does. Second, every winning move provides structure to the LLM upfront — the test intent, the refactor category, the schema — rather than asking the model to infer it. These two themes recur through Section 12's templates and Section 13's honest-limits discussion.

Apply to DV
  • Audit your top 10 prompts against the seven patterns. The patterns missing are usually the lowest-hanging adoption fruit.
  • For each adopted pattern, instrument a metric so you can measure the lift on your eval suite over time. The SWE results are large; the DV results should be too.

11. HW-LLM Frameworks Already in the Wild

The 2024-2026 hardware-LLM research has converged on a small number of named systems. Each one reveals a specific lesson about prompting for HDL/UVM that you can borrow without adopting the whole system.

Research Ten published HW-LLM systems, with the lesson each contributes: HAVEN (arxiv 2604.27643) plan-then-template, 100% compile, 90.6% coverage. UVM² (arxiv 2504.19959) domain-knowledge prompts + syntactic constraints + iterative refinement. UVMarvel (arxiv 2605.04704) multi-agent per protocol, 95.65% coverage. VeriGRAG (arxiv 2510.15914) structure-aware soft prompts. FVDebug (arxiv 2510.15906) causal graph synthesis + agentic exploration for waveform/RTL/spec debug. MEIC (arxiv 2405.06840) iterative debug with bounded progress per turn. LASHED (arxiv 2504.21770) LLM + static analysis for early RTL bug detection. Self-HWDebug (arxiv 2405.12347) self-instructing debug from vulnerable/secure RTL pairs. ReasoningV (arxiv 2504.14560) reasoning specialization for Verilog. Proof2Silicon (arxiv 2509.06239) RL prompt repair from formal feedback. CVDP (arxiv 2506.14074) benchmark establishes the 34% pass@1 ceiling.
  1. Never let the LLM emit HDL directly without scaffolding. This is the single lesson all 10 systems share. HAVEN's split (LLM → plan JSON; template engine → UVM) is the canonical pattern; the 100% compile success and 90.6% coverage numbers are the proof.
  2. Iterate with feedback, not in isolation. UVM², MEIC, and Proof2Silicon all use iterative refinement against a verifier signal (coverage, syntax, formal property). Single-shot generation is what produces the disappointing CVDP numbers; iteration is what closes the gap.
  3. Combine LLM with classical tools. LASHED pairs LLM with static analysis. Proof2Silicon pairs LLM with formal verification. The hybrid systems beat LLM-alone systems consistently.
  4. Specialize agents per protocol or subsystem. UVMarvel uses a different agent per bus protocol and reaches 95.65% coverage. Generalist agents over-fit to common protocols and under-perform on rarer ones.
  5. Borrow the patterns without adopting the systems. You do not need to deploy HAVEN to use its plan-then-template idea in your own prompts. The published systems are demonstrations; the patterns are reusable.
Apply to DV
  • Pick one HW-LLM lesson and apply it to your current workflow this month. Plan-then-template is the highest-leverage starter.
  • Track your own pass@1 number on a small local benchmark. The 34% CVDP ceiling is a public-research average; your number with structured prompts and iteration should be substantially higher on tasks scoped to your IP family.
  • When evaluating a vendor's AI-DV tool, ask which of the 10 patterns above it implements and which it skips. The honest vendors can answer.

12. A Concrete DV Prompt Template Library

Six ready-to-paste prompt templates, each annotated with the research pattern it borrows from. Replace the {{variables}} with your specifics. None of these templates are theoretical; each composes 2-3 of the techniques from Sections 1-10.

Template 1 — RTL Module from Spec (decomposition + few-shot + intent-first)

[SYSTEM]
You are a senior RTL designer writing SystemVerilog for a verification
team to validate. Match the style of the provided examples. Never invent
functionality not present in the spec.

[USER]
Generate a SystemVerilog module implementing the feature below.

Decomposition (handle in order):
1. Identify all signals required from the spec
2. Identify the FSM states (if any)
3. Identify the data flow (combinational vs sequential)
4. Generate the port list
5. Generate the module body
6. State any assumptions in a final comment

Spec section:
{{spec_excerpt}}

Module name: {{module_name}}
Style examples (from our codebase):
{{example_1}}
{{example_2}}

Why this works: The numbered decomposition (Section 5) prevents the model from emitting a monolithic blob. The codebase-derived examples (Section 6) carry team idioms the model would not infer. The explicit "never invent" instruction caps the hallucination surface.

Template 2 — UVM Agent Scaffolding (role + decomposition + few-shot + scope limits)

[SYSTEM]
You are a senior UVM verification engineer scaffolding a new agent.
Match the team's existing patterns. Never invent functionality not
explicitly asked for. Default to non-blocking, event-driven design.

[USER]
Generate UVM agent scaffolding for the interface below.

Interface definition:
{{interface_signature}}

Reference monitor (the team's style guide):
{{reference_monitor_code}}

Output, in order:
1. Transaction class (sequence_item) with rand fields and UUID stamping
2. Driver class with run_phase
3. Monitor class with run_phase and analysis_port
4. Sequencer typedef
5. Agent class with build_phase + connect_phase

Do NOT generate:
- The sequence library (separate prompt)
- The scoreboard (separate prompt)
- Any test class

Why this works: Role priming (Section 1) sets the assistant voice. Decomposition with explicit ordering (Section 5) keeps the agent components consistent. The "do NOT" list prevents the helpful-assistant scope creep that breaks generated TBs.

Template 3 — Failure RCA (Hypothesis-Rank) (CoT + multi-gen self-consistency + constrained-output JSON)

[SYSTEM]
You are a senior DV engineer triaging a UVM regression failure. Your
job is ranked hypotheses with falsifying experiments, not free-form
analysis. Reason from evidence in the bundle, never from generic
knowledge of the protocol.

[USER]
Analyze this failure bundle and propose the 3 most likely root causes
ranked by probability. For each cause:
- one-sentence explanation grounded in a specific bundle field
- ONE falsifying experiment (plusarg, variant, probe) runnable in <5 min
- the bundle field (line/event/signal) that supports the hypothesis

Return JSON only, no prose:
[
  {"cause": "...", "prob": 0.XX, "falsifier": "...", "evidence_field": "..."},
  ...
]

Bundle:
{{bundle_json}}

Why this works: Constrained-decoding output (Section 8) forces the structure to be machine-checkable. Forcing "grounded in a specific bundle field" closes the hallucination escape hatch. Run the same prompt with temperature > 0 three times for self-consistency (Section 3); accept the hypothesis appearing as #1 across all three runs.

Template 4 — Coverage Plan from Spec (decomposition + JSON schema + intent-first)

[SYSTEM]
You are a UVM verification engineer producing a coverage plan. Output
is JSON consumed by a covergroup generator. Cover features, not signals.
Every cross must have a reason.

[USER]
Generate a coverage plan from the spec section below.

Spec section:
{{spec_section}}

Decompose into:
1. Feature categories (top-level groups)
2. Per-feature bins (specific values to cover)
3. Per-feature crosses (combinations - each must have a reason)
4. Per-feature negative scenarios (error and corner cases)

Output JSON shape:
{
  "features": [
    {
      "name": "...",
      "bins": [{"name": "...", "values": [...]}],
      "crosses": [{"name": "...", "with": [...], "reason": "..."}],
      "negatives": [{"name": "...", "scenario": "..."}]
    }
  ]
}

Why this works: "Cover features, not signals" is intent-first prompting (Section 10 pattern 3) at the system level. The "every cross must have a reason" constraint prevents combinatorial blowup of meaningless crosses. JSON shape pinned for downstream tooling.

Template 5 — NL to SVA Translation (structured-reasoning template + bounded output)

[SYSTEM]
You are a formal verification engineer translating natural-language
requirements into SystemVerilog Assertions. Output must be syntactically
valid SVA. Use vacuous-pass-aware patterns. Use semi-formal reasoning:
state premises, trace execution, derive the property.

[USER]
Translate this requirement into one SVA property.

Requirement:
"{{nl_requirement}}"

Available signals:
{{signal_list_with_widths}}

Clock: clk (posedge), Reset: rst_n (active low)

Reasoning structure:
1. Premises: what signals participate and what their roles are
2. Execution: cycle-by-cycle trace of the requirement
3. Derivation: the SVA expression

Output:
- property_name: ...
- property body: ... (assert property)
- cover property: ... (proves the property is reachable, not vacuously true)
- non-vacuity note: one sentence explaining why this property cannot
  be vacuously true

Why this works: Meta's semi-formal reasoning template (Section 10 pattern 4) is the direct ancestor — premises, execution, derivation. The explicit non-vacuity check addresses the most common SVA bug: properties that pass because nothing triggers them.

Template 6 — UVM Refactor (Subcategory-First) (subcategory-explanation pattern + role + scope discipline)

[SYSTEM]
You are a senior UVM engineer refactoring legacy testbench code.
You must FIRST identify the refactoring category, THEN propose the
change. Never refactor without naming the category from the list below.

[USER]
Refactor this UVM code. The refactoring category is: {{category}}.

Categories available:
- extract_base_class: pull common functionality into a base class
- replace_inheritance_with_composition: convert is-a to has-a
- virtual_sequence_extraction: factor sequences from a monolithic test
- agent_split: separate driver/monitor/sequencer responsibilities
- factory_override_simplification: collapse N overrides into config_db
- config_object_introduction: replace string parameters with typed config

Code to refactor:
{{legacy_code}}

Output:
1. Confirm the category and explain in 2 sentences why it applies
2. The refactored code
3. Backwards-compatibility note: what tests or sequences need to change
4. Risk assessment: 1-10, with one sentence justification

Why this works: The subcategory-explanation pattern (Section 10 pattern 7) is what took refactoring success from 15.6% to 86.7% in the SWE literature. The explicit refactoring categories prevent the model from inventing its own. The risk-assessment line forces the model to consider what it is changing — not just produce the change.

Apply to DV
  • Adopt one template this week. The hypothesis-rank template (Template 3) has the highest leverage for engineers actively triaging regressions.
  • Version templates in your team repo. Add them to the eval suite from Section 9 so model upgrades cannot silently break them.
  • Build one new template per quarter as you discover prompts that consistently work for your team's workflow.

13. What Does Not Work in DV (Honest)

An honest accounting of where the research and the lived experience agree the LLM falls short. Knowing these in advance is what keeps you from wasting an afternoon on a prompt that was never going to work.

Research The 2026 RTL survey explicitly names naive CoT as ineffective for IC design. CVDP (arxiv 2506.14074) shows SOTA models hit only 34% pass@1 on hardware-specific tasks. ProtocolLLM (arxiv 2506.07945) shows even syntactic validity is unreliable on novel SystemVerilog testbench code. The effective-context-window paper (arxiv 2509.21361) shows degradation past ~1,000 tokens of operational context even for frontier models with advertised 200K windows.
  1. Naive chain-of-thought for RTL design. "Let's think step by step" produces generic reasoning that does not map to the actual structure of HDL synthesis. The fix is HW-aware CoT scaffolding (Section 2), not abandoning CoT entirely.
  2. Long raw log dumps as context. Pasting 2,000 lines of UVM log into a chat is the textbook trigger of the effective-window degradation. Use the JSON bundle pattern instead — 200 structured events beat 2,000 lines of prose every time.
  3. Asking for "novel" protocol implementations. The model knows AXI, PCIe, USB, and AHB because the training data does. Your team's proprietary or pre-release protocols are not in there; generated code looks right and is wrong.
  4. Trusting LLM-emitted line numbers. Frontier models routinely produce line numbers off by 4-10 lines, even for code in the prompt. Treat the line number as a neighborhood pointer, never as an exact location. The fix is to verify before applying.
  5. Random few-shot examples. The research is unambiguous: random or text-similarity examples are sub-optimal. If you do not invest in example selection (Section 6), you leave significant performance on the table.
  6. One-shot complex IP verification. "Verify this PCIe controller" as a single prompt produces vague boilerplate. Decomposition is non-negotiable for tasks larger than a single component.
  7. Believing published pass@1 numbers as your own ceiling. Benchmarks like CVDP report averages across heterogeneous tasks. Your number on tasks scoped to your IP family with structured prompts and iteration is almost certainly higher — measure it yourself.
  8. Single-sample answers for high-stakes decisions. Any output that goes into shipping silicon (an SVA the design relies on, a coverage point that gates signoff, an RTL fix) should pass either self-consistency, formal verification, or human review — not all three is fine, none of three is not.
Apply to DV
  • Print this section as a poster. The bottom three failure modes (LLM line numbers, one-shot IP verification, single-sample high-stakes) are the ones a junior engineer is most likely to learn the hard way.
  • For tasks the model genuinely cannot do (novel protocols, multi-cycle deep reasoning), stop trying. Fall back to traditional methods; do not waste your morning on iteration #4 of a prompt that was never going to converge.

Reading List

Prompting Foundations & Surveys

  • arxiv 2406.06608 — The Prompt Report (PRISMA-grounded survey of 58 techniques)
  • arxiv 2402.07927 — Systematic Survey of Prompt Engineering
  • arxiv 2407.12994 — Prompt Engineering Methods for NLP Tasks
  • Liu et al. 2026, Frontiers of CS — Comprehensive Taxonomy of Prompt Engineering Techniques

Reasoning (CoT, Self-Consistency, Tree-of-Thoughts)

  • arxiv 2201.11903 — Wei et al., Chain-of-Thought Prompting
  • arxiv 2203.11171 — Wang et al., Self-Consistency Improves CoT
  • arxiv 2305.10601 — Yao et al., Tree-of-Thoughts
  • arxiv 2401.14295 — Besta et al., Chains, Trees, Graphs of Thoughts
  • arxiv 2510.01069 — Typed CoT / Certified Self-Consistency (2026)
  • arxiv 2603.08999 — Confidence-Aware Self-Consistency (2026)

Decomposition & Few-Shot

  • arxiv 2205.10625 — Zhou et al., Least-to-Most Prompting
  • arxiv 2210.02406 — Khot et al., Decomposed Prompting
  • arxiv 2310.09748 — LAIL: LLM-Aware ICL for Code Generation
  • arxiv 2305.14210 — Skill-Based Few-Shot Selection
  • arxiv 2412.02906 — Does Few-Shot Help LLM Code Synthesis?

Agentic Patterns

  • arxiv 2210.03629 — Yao et al., ReAct
  • arxiv 2303.11366 — Shinn et al., Reflexion
  • arxiv 2509.06239 — Proof2Silicon: RL Prompt Repair from Formal Feedback

Constrained Decoding

  • arxiv 2603.03305 — Draft-Conditioned Constrained Decoding
  • llguidance (github.com/guidance-ai/llguidance) — ~50µs/token CFG enforcement
  • Outlines library — JSON-schema-compiled valid-token lookup

Meta-Prompting & Promptware Engineering

  • arxiv 2502.00728 — Meta-Prompt Optimization for Sequential Decision Making
  • arxiv 2503.02400 — Promptware Engineering

SWE Prompting Empirical Studies

  • arxiv 2504.06939 — FeedbackEval (code repair)
  • arxiv 2603.19399 — DePro (debug)
  • arxiv 2506.13186 — Empirical Evaluation of APR
  • arxiv 2505.04441 — Execution Traces for Program Repair
  • arxiv 2506.09002 — PALM (Rust unit test coverage)
  • arxiv 2407.00225 — Prompt Engineering for Unit Test Generation
  • arxiv 2402.00097 — Code-Aware Prompting for Coverage-Guided Tests
  • arxiv 2508.14419 — Static Analysis as Feedback Loop
  • arxiv 2411.02320 — Empirical Study on Code Refactoring (15.6% → 86.7%)
  • arxiv 2303.07839 — ChatGPT Prompt Patterns for Code Quality

HW-LLM Frameworks & Benchmarks

  • arxiv 2604.27643 — HAVEN (UVM TB synthesis, 100% compile, 90.6% coverage)
  • arxiv 2504.19959 — UVM² (LLM-aided UVM machine)
  • arxiv 2605.04704 — UVMarvel (subsystem-level, 95.65% coverage)
  • arxiv 2510.15914 — VeriGRAG (structure-aware soft prompts)
  • arxiv 2510.15906 — FVDebug (waveform + RTL + spec debug)
  • arxiv 2405.06840 — MEIC (iterative RTL debug)
  • arxiv 2504.21770 — LASHED (LLM + static analysis)
  • arxiv 2405.12347 — Self-HWDebug (self-instructing security debug)
  • arxiv 2504.14560 — ReasoningV (efficient Verilog generation)
  • arxiv 2506.14074 — CVDP benchmark (34% SOTA pass@1)
  • arxiv 2506.07945 — ProtocolLLM (SV testbench benchmark)
  • arxiv 2212.11140 — Benchmarking LLMs for Verilog (foundational)

Context Engineering

  • Anthropic, Effective Context Engineering for AI Agents (2025) — the canonical industry essay
  • arxiv 2509.21361 — Maximum Effective Context Window
  • arxiv 2603.04814 — Beyond the Context Window (fact-memory vs long-context)
Author
Milan Kubavat
Sharing knowledge about silicon verification, hardware design, and engineering insights.

Comments (0)

Leave a Comment