Context Engineering for DV: From Failure Bundles to Agent Memory

In 2025, an estimated 65% of enterprise AI failures were attributed to context drift or memory loss during multi-step reasoning. That is not an edge case — it is the dominant failure mode. The field has responded by reorganizing around context engineering: the discipline of curating what tokens reach the model at each step, what the model remembers across steps, and how it retrieves what it needs without drowning in what it does not.

This post is the deep-dive on what that means for Design Verification. Eleven sections cover the effective-context-window research, the context taxonomy, minimum-viable-context discipline, memory systems for long-running DV agents, compression techniques, tool design, prompt caching economics, structured output as implicit context, the bundle anatomy that makes DV-specific context work, agent memory for long-running workflows, and the honest limits of all of it.

The structure follows the same pattern as the rest of this AI series — each section opens with a research callout citing the specific paper, numbered playbook patterns provide the takeaway, and DV-application blocks ground each technique in your UVM testbench or RTL workflow. Cross-links to the AI Playbook, DV Prompts, Full Pipeline, and Structured Logging posts thread throughout where appropriate.

1. The Effective Context Window Crisis

The single most actionable finding the DV community has not internalized: advertised context windows are not effective context windows. The gap is large, the failure modes are predictable, and pretending otherwise is the #1 source of disappointing LLM-debug experiences.

Research arxiv 2509.21361 Context Is What You Need: The Maximum Effective Context Window for Real World Limits of LLMs — effective capacity is typically 60-70% of advertised; some tasks degrade by up to 99%. The Context Rot analysis (Morph LLM, 2026) names three compounding degradation mechanisms. The lost-in-the-middle phenomenon, first documented by Stanford/UC Berkeley researchers in 2023 and refined repeatedly since, shows 30%+ accuracy drops for information positioned in the middle of context vs. the start or end. The 2026 enterprise-AI failure analyses attribute the majority of multi-step reasoning failures to context drift.
  1. Measure your model's real MECW on your workload. Do not trust the marketing number. Build a 10-prompt eval set with varying context sizes (1K, 5K, 20K, 50K, 100K tokens) and measure accuracy degradation. The curve is what tells you where to cap.
  2. Place critical information at the start and end. The lost-in-the-middle research is unambiguous: middle-positioned facts are functionally ignored at large context sizes. The failing event in a bundle goes at the top; the system instructions stay at the top of the system prompt; long mid-prompt context dumps are anti-patterns.
  3. Cap at 50-60% of the advertised window. If your model claims 200K tokens, treat 120K as the soft ceiling and 60K as the comfortable working zone. The marginal token past 60% is doing less work than the first one.
  4. Avoid mid-context noise. Distractor interference — semantically similar but irrelevant content — is a measurable third degradation mechanism. The fix is not to include “helpful background” that is not load-bearing for the task.
  5. Re-measure on every model upgrade. MECW changes when the underlying model changes. The eval set from step 1 is your regression suite; run it on every provider update.
Apply to DV
  • Never paste raw regression logs into a chat. The 2,000-line log dump is the textbook trigger of all three degradation mechanisms.
  • In your failure bundles (Section 9), the failure event lives at the top of the bundle JSON, before the 200-event context window. The agent sees the punchline first.
  • If your team uses a 200K model, set the bundle composer's max_tokens at 60-80K, not 200K. The compose-and-pray approach kills accuracy long before it kills cost.

2. Context Taxonomy: What Lives in the Window

Not all context is the same. A clean taxonomy of what types of information compete for the window lets you optimize each independently — cache the stable, prune the ephemeral, retrieve the on-demand.

Research The Anthropic Effective Context Engineering for AI Agents essay (September 2025) and the LangChain context-engineering writeup converge on a six-part taxonomy. arxiv 2603.09619 Context Engineering: From Prompts to Corporate Multi-Agent Architecture formalizes the lifecycle distinction: stable, semi-stable, and ephemeral context have different caching, retrieval, and pruning strategies.
  1. System instructions. Stable across the session. Cache hot. This is your prompt template, role priming, output schema. Changes only on team-level prompt evolution.
  2. In-context examples (few-shot). Semi-stable. Swap the examples by task family (AXI debug, PCIe debug, scoreboard generation). Cache by task family; refresh when the team adds canonical examples.
  3. Retrieved knowledge. Ephemeral. Just-in-time. RTL excerpts, spec sections, past-bug context come in via tool calls when the agent asks for them — never preemptively loaded.
  4. Tool outputs. Ephemeral. Summarize aggressively and drop the raw form. The 800-line log slice from a query_log call should be reduced to 30 key events before the next reasoning step.
  5. Conversation history. Semi-stable but compounds dangerously. Prune aggressively. At session checkpoints (every 8-10 turns), re-summarize the conversation into a 200-token state synopsis and let the rest fall off.
  6. Scratchpad / think tool. Ephemeral. Fresh per agent step. The reasoning output is not preserved across iterations — only the conclusions feed forward.
Apply to DV
  • System instructions: your senior-DV-engineer role priming, the bundle schema description, the output JSON shape. Cache hot.
  • Examples: keep a per-protocol example library (axi, pcie, usb, custom). The bundle composer picks the relevant examples based on the failure's protocol family.
  • Retrieved: RTL excerpts via get_rtl_excerpt, past bugs via query_fingerprint_db, design constraints via query_spec. None loaded upfront.
  • Tool outputs: run_smoke returns pass/fail + log digest, not the full log. The agent does not need 50K tokens of new log to learn the smoke passed.

3. The Minimum-Viable-Context Principle

The 2025-2026 industry consensus has shifted from “more tokens equal more capability” to just-in-time retrieval and progressive disclosure. The agent should fetch what it needs when it needs it, not start the session with everything that might be relevant.

Research Anthropic's Effective Context Engineering for AI Agents (September 2025) is the canonical industry essay and is explicit about just-in-time retrieval as the recommended strategy for long-running agents. Claude Code is the canonical implementation: maintain lightweight identifiers (file paths, line ranges, IDs) and let tools like grep/glob/read dynamically load only what the agent requests. The State of Context Engineering 2026 (Aurimas Griciunas) and the LogRocket LLM context problem analysis (2026) corroborate: progressive disclosure beats upfront loading on every measured workload.
  1. Lightweight references over data dumps. The bundle composer should emit file_paths, signal_names, fingerprint_ids, commit_hashes — not the full contents. The agent fetches contents via tools when it decides the reference is relevant.
  2. Tools for on-demand load. get_rtl_excerpt(file, line_range), query_log(filter, max_events), get_commit_diff(hash). Each tool is the equivalent of grep/glob/read for the DV domain.
  3. Progressive disclosure: broad to narrow. The first tool call returns a summary (e.g., 5 candidate RTL files); subsequent calls drill into the one the agent picks. Avoid round-trips on broad scans.
  4. Verify the agent did not over-ask. Cap tool-call budgets per session (8-12 calls maximum). When the agent is over budget, that is your signal the task scope is wrong — not that you should raise the budget.
Apply to DV
  • Your failure bundle (Section 9) carries identifiers, not data: failing transaction txn_id, suspect signal names, candidate RTL file paths, fingerprint hash, run header. The agent calls tools to fetch what it actually needs.
  • For SoC-scale debug, the first tool call returns the per-IP failure summary; the agent picks the IP; subsequent calls drill into that IP only. Never pre-load all IPs.
  • Cap your debug agent at 10 tool calls per session. If it needs more, the bundle was wrong or the task was too big — not the agent's fault, but yours.

4. Memory Systems for DV Agents

Memory has become a first-class architectural component — not an afterthought of the model's context window. The 2024-2026 research has settled on a four-type taxonomy borrowed from cognitive science, and the DV-side mapping is direct once you see it.

Research arxiv 2502.06975 Episodic Memory is the Missing Piece for Long-Term LLM Agents — the foundational argument. arxiv 2604.04853 MemMachine: A Ground-Truth-Preserving Memory System for Personalized AI Agents (2026) is the canonical implementation. The 2026 State of AI Agent Memory report (Mem0) and the GitHub Memory in the Age of AI Agents survey converge on four memory types. Recent 2026 papers (MemRL, Agentic Memory) extend with self-evolving and unified short-long memory management.
  1. Short-term memory (STM) is the context window itself. Do not fight it — treat it as scarce. Apply Sections 1-3 ruthlessly: cap context, prune mid-window noise, use just-in-time retrieval.
  2. Episodic memory: specific past experiences. The vector database of past failure bundles + their resolutions. When a new failure arrives, semantic-similarity-search the episodic store to find “we have seen something like this before.” The bundle composer surfaces matches as fingerprint_siblings.
  3. Semantic memory: distilled facts and constraints. Protocol invariants (AXI burst constraints, PCIe LTSSM rules), design constraints (clock-domain mappings, power-state transitions), team conventions (UVM agent patterns, naming standards). Updated rarely; consulted often.
  4. Procedural memory: learned strategies. Debug recipes that worked. “For SCB_MISMATCH fingerprints involving D3 power state, check power-controller reset domain first.” Built from the postmortem stage of the Full Pipeline; consulted by the bundle composer at agent invocation.
  5. Memory consolidation: episodic to semantic over time. The MemMachine pattern: episodic stores raw experience; a periodic consolidation pass distills recurring patterns into semantic facts. For DV: monthly job that converts “these 47 past bugs all involved CDC issues on clk_b” into a semantic-memory entry “clk_b CDC is a frequent failure class — check first.”
Apply to DV
  • Stand up an episodic memory store (SQLite + vector index): one row per resolved fingerprint, embedding of the bundle, link to the fix commit. The Full Pipeline post's Stage 6 already writes this; this is its purpose.
  • Stand up a semantic memory store: a structured YAML or JSON file per IP capturing “known invariants” (signal width assumptions, FSM transition rules, power-state semantics). Updated on spec change; consulted on every debug session.
  • Stand up a procedural memory store: a runbook indexed by failure-class fingerprint — “for fingerprint class X, run plusarg Y first.” Build this empirically from successful debug sessions.
  • Memory consolidation: a quarterly review that asks “what episodic entries clustered into recurring patterns?” and writes them into semantic memory. This is the maintenance step most teams skip and most regret.

5. Context Compression Techniques

Compression is the engineering response to the effective-window crisis: keep the load-bearing tokens, shed the rest. The 2026 research has converged on three concrete approaches, each with measured token-reduction numbers in the 26-54% range without task-success degradation.

Research arxiv 2601.07190 Active Context Compression: Autonomous Memory Management in LLM Agents (Focus architecture, slime-mold-inspired exploration) — addresses “context bloat” in long-horizon SWE tasks by autonomous pruning. arxiv 2510.00615 ACON: Optimizing Context Compression for Long-Horizon LLM Agents — failure-driven, task-aware compression guideline optimization, reducing peak tokens by 26-54% while preserving task success. arxiv 2510.08907 Autoencoding-Free Context Compression via Contextual Semantic Anchors (SAC) consistently outperforms prior compression methods.
  1. Anchored iterative summarization. At checkpoints, summarize prior turns into a 200-token state synopsis anchored to key entities (failure event, hypothesis, last tool result). Subsequent turns reference the synopsis, not the raw history.
  2. Failure-driven guideline optimization (ACON). When the agent fails to converge, the meta-loop asks “what part of the context was load-bearing? What was noise?” and updates the compression policy for next time. Compression learns from misuse.
  3. Autonomous pruning (Focus). Agents decide for themselves when to consolidate raw history into a persistent “knowledge block” and drop the rest. The slime-mold metaphor: keep the trails to food, prune the dead ends.
  4. Semantic anchor compression (SAC). Instead of summarizing prose, compress around named entities (the failing signal, the suspect FSM, the relevant txn_id) and let the model reconstruct details from anchors at inference time.
  5. Provider-native compaction. Anthropic and OpenAI both offer compaction APIs that summarize prior turns automatically. Cheaper than DIY; less control. Use when your eval suite says quality is preserved.
Apply to DV
  • For long debug sessions: at every 5 tool calls, re-summarize the session into a state synopsis — failing event, current hypothesis, what has been falsified, what is unknown. The next 5 calls operate on the synopsis, not the raw trace.
  • Compress regression history into per-fingerprint summaries rather than carrying every past failure. A summary line per fingerprint with hit-count and last-resolution outperforms the raw event stream for context purposes.
  • Use semantic anchors for DV: the txn_id is the natural anchor for a transaction; the fingerprint is the natural anchor for a failure class; the signal name is the natural anchor for a debug investigation. Compress around these.

6. Tool Design as Context Discipline

Tool design is context engineering by other means. A well-designed tool returns minimum-viable context; a poorly-designed tool floods the window with noise and burns the budget you spent the previous five sections protecting. Anthropic's 2025-2026 guidance is unusually concrete.

Research Anthropic Writing Effective Tools for AI Agents (2025) — core principles: tools need different design than human-facing APIs because they must be protected against the LLM “chaos monkey”; choose high-leverage tools; clear distinct names; human-readable fields beat raw IDs; pagination, truncation, filtering as defaults. arxiv 2505.18135 Tool Preferences in Agentic LLMs are Unreliable — agents under-use even well-designed tools, over-use poorly-named ones. arxiv 2412.04093 Practical Considerations for Agentic LLM Systems covers token-aware tool design at scale.
  1. High-leverage tools, not thin wrappers. A tool that expands agent capability is worth a tool slot; a tool that just relays an API call is not. query_log(filter, max_events) with a focused filter beats read_file(path) applied to a 50MB log.
  2. Clear, distinct names. No two tools should be plausibly confusable. get_rtl_excerpt vs read_rtl_file is exactly the confusion to avoid — pick one verb, document the contract, deprecate the other.
  3. Human-readable fields beat raw IDs. Return {"signal": "axi_wb_inflight", "value": 0, "at_ns": 4960}, not {"s": 42, "v": 0, "t": 4960}. The agent reasons about names, not opaque integers.
  4. Pagination, truncation, filtering as defaults. Every tool should return at most a few hundred tokens by default. A tool that could return 50K tokens must have explicit pagination parameters and the agent must opt in.
  5. Defensive against misuse. Validate parameters at the tool boundary; return clear error messages the model can recover from. The agent will pass nonsense; your tool should reject it gracefully without filling the window with stack traces.
Apply to DV
  • The DV agent tool catalog: query_log(filter, max_events=200), get_rtl_excerpt(file, line, context_lines=20), run_smoke(test, seed, plusargs, timeout=300), query_fingerprint_db(fingerprint), get_commit_diff(hash), list_signal_history(signal, time_range). Six tools, each with explicit limits.
  • Tool outputs are digests, not raw artifacts. run_smoke returns {pass: true/false, duration_s, error_class, last_10_events}, not the full 50K-line simulation log.
  • If a tool has more than 5 parameters, split it. If a tool's name needs a paragraph to explain, rename it.

7. Prompt Caching Economics for DV Loops

Prompt caching is the highest-ROI cost lever for any agentic workload that re-queries the same context dozens of times per session. UVM debug loops are exactly that workload. For teams running this at scale, not using prompt caching is not a cost optimization to skip — it is unaffordable not to use.

Research Anthropic prompt caching: writes cost 1.25x normal input rate; reads cost 10% of normal input rate (a 90% discount). OpenAI's cached prefix is billed at 50% of normal. Real-world reports show 70-90% token cost reduction in agent loops; Anthropic claims up to 85% latency reduction for long prompts with TTFT dropping from 5s to under 200ms. arxiv 2601.06007 Don't Break the Cache: An Evaluation of Prompt Caching for Long-Horizon Agentic Tasks — cache hit rate becomes the dominant cost factor at scale.
  1. Order context for cacheability. Stable content first, volatile content last. The system prompt, role priming, output schema, and codebase examples go at the top — they do not change between requests in the session. The failure event and per-request bundle slice go at the bottom — they change every call.
  2. Cache breakpoint design (Anthropic). Anthropic gives you explicit cache breakpoints with configurable TTL (5 minutes default). Place breakpoints after each stable block. The agent gets a 10% read on everything before the breakpoint, full price only on the new tokens after.
  3. Cache invalidation strategy. When does the cache expire? When the system prompt changes, when the schema version bumps, when the team examples are refreshed. Plan invalidation around team release cadence, not silently.
  4. Track cache hit rate as a first-class metric. If your agent's cache hit rate is below 60%, your context ordering is wrong. If it is above 90%, you are doing context engineering right.
Apply to DV
  • For UVM debug sessions where the agent re-queries the same bundle context across 5-15 ReAct iterations, prompt caching cuts per-iteration cost by ~90% on cached tokens. A $0.30 debug session becomes a $0.05 debug session at no quality cost.
  • Order your debug agent context: [stable: system prompt + role + bundle schema + protocol examples] — cache breakpoint — [volatile: failure event + 200 context events + tool outputs from this turn].
  • Add cache-hit-rate to your debug-pipeline dashboard. If you do not measure it, the cache is silently degrading and you will discover only when the bill arrives.

8. Structured Output as Implicit Context

Schema-enforced structured output is not just a parsing convenience — it is implicit context. The schema you pin to the response tells the model exactly what fields to produce, in what order, with what types. That schema is doing context-engineering work without consuming context tokens.

Research The 2026 industry consensus identifies three reliability levels: prompt engineering for structured output (80-95% valid), function calling/tool use (95-99%), and native structured output with constrained decoding (100% schema-valid by construction). OpenAI Strict Mode (August 2024), Anthropic native structured output (late 2025), and Google Gemini equivalent capability all use finite-state-machine compilation from JSON Schema. XGrammar (March 2026) is the default structured-generation backend for vLLM, SGLang, and TensorRT-LLM at <40 microseconds per token. Outlines is the open-source Python library pioneering grammar-based generation.
  1. Three reliability levels — know which one you are using. If your prompt says “return JSON with fields X, Y, Z” you are at level 1 (80-95%). If you use the provider's tool-call interface, you are at level 2 (95-99%). If you use native structured output with FSM enforcement, you are at level 3 (100% schema-valid).
  2. Schema-as-context. The schema describes the output shape and the model treats it as instruction. A well-named schema field (fingerprint_hash) is more useful than an instruction-prose equivalent (“include a fingerprint hash here”).
  3. Schema versioning is part of context engineering. When you change a schema, you change the implicit context. Version your schemas; include the version in cache invalidation; include it in your prompt eval suite.
  4. Validate even with constrained decoding. Schema-valid does not mean semantically correct. A schema-valid coverage plan can still be a bad coverage plan. Constrained decoding eliminates one class of errors; downstream validation handles the rest.
Apply to DV
  • Maintain a DV schema library: hypothesis_rank.schema.json, coverage_plan.schema.json, fix_proposal.schema.json, refactor_proposal.schema.json. Version each.
  • Use level-3 native structured output for any agent response that feeds downstream tooling. Use level-1 only when humans are the only consumer.
  • The schema definitions themselves are part of your team's context-engineering surface area. Treat them with the same review discipline as production code.

9. DV-Specific Context Patterns: Bundle Anatomy

Everything above sets up this section. The DV-specific context pattern — the failure bundle — is what makes context engineering land for hardware verification. This is the deep-dive on bundle anatomy, RTL excerpt selection, signal slice formatting, DUT state representation, and the multi-IP coordination patterns that emerge at SoC scale.

Research The Anthropic context-engineering essay applied to coding agents (Claude Code) is the architectural template; the DV adaptation builds on the Structured Logging JSONL substrate and the Full Pipeline bundle pattern. The 2026 industry coverage of Siemens' agentic verification toolkit shows MCP-based context exposure: real-time design and verification state available to agents through a standardized protocol. arxiv 2512.06247 DUET: Agentic Design Understanding via Experimentation and Testing demonstrates agentic dynamic exploration of RTL with the agent generating its own experiments to build context.
  1. Bundle anatomy — six well-typed sections. The DV failure bundle is a JSON document with: (a) failure — the triggering event, (b) context_events — ring-buffer dump of pre-failure events (Structured Logging Pattern 9), (c) rtl_excerpts — relevant RTL code snippets with reasons for inclusion, (d) fingerprint_siblings — past bugs with the same fingerprint plus their resolutions, (e) recent_commits — RTL git log for the last N days scoped to the failing IP, (f) run_header — seed, plusargs, tool version, RTL hash for replay.
  2. RTL excerpt selection — symbol-indexed, not text-grep. Build a Verible / sv-parser / Surelog index of your RTL once, cached. For each failure, the bundle composer queries the index by signal names appearing in the context events, scoped to the failing IP. Returns the file path, line range, and 25 surrounding lines for each match, ranked by hit count. Never blind grep across the SoC.
  3. Signal slices as DUT_STATE events. Your bind-based harness (see Debug page Pattern 11) emits DUT_STATE JSONL events on key triggers: FSM transitions, queue depth changes, register writes. These appear as first-class events in the context_events stream — the agent sees them inline, not as a separate waveform query.
  4. DUT state representation. A snapshot DUT_STATE event contains the current state of the things that change rarely but matter critically: FSM states, queue depths, key configuration registers, power state, clock domain, mode/profile. Format as flat key-value pairs in the JSONL event; the agent reads them with zero ceremony.
  5. Per-IP bundle scoping for SoC debug. At SoC scale, never compose a bundle that spans IPs by default. The first-pass bundle is per-IP; cross-IP correlation happens via shared txn_id threading (Structured Logging Pattern 3). The agent decides when it needs to cross IP boundaries and fetches additional per-IP bundles via tool calls.
  6. MCP-based context exposure (Siemens pattern). The Model Context Protocol gives you a standardized server interface for exposing DV tools to any MCP-compatible client (Claude Code, IDE assistants, custom agents). Each tool from Section 6 becomes an MCP endpoint; the protocol handles the rest.
  7. Bundle hash for deterministic replay. Every bundle gets a content hash. Cache prompt+response by hash. The same failure replays to the same agent response, every time. Reproducibility is non-negotiable when the agent's output goes anywhere near production silicon decisions.
Apply to DV
  • If you build only one piece of context-engineering infrastructure this quarter, build the bundle composer. Everything else — agent design, memory, prompt caching — assumes the bundle exists.
  • Adopt MCP early. The 2026 vendor landscape (Siemens, Cadence experimental, Synopsys announced) is converging on MCP as the integration protocol. Build your DV tool catalog as MCP-compatible from day one.
  • The bundle composer is the single most important piece of Python in your DV-AI stack. Treat it accordingly: code review, unit tests, schema versioning, performance budget.
  • For multi-team SoC environments, agree on a shared bundle schema across teams. The schema is a contract; each IP team produces bundles to it, agents consume bundles in the same shape regardless of source.

10. Agent Memory for Long-Running DV Workflows

Section 4 laid out the memory taxonomy. This section grounds it in the long-running workflows DV teams actually run: regression seasons that span weeks, projects that span quarters, IP families that span years. Memory is what turns the agent from a smart one-shot tool into a teammate that gets better with the project.

Research arxiv 2604.04853 MemMachine — the canonical ground-truth-preserving memory system architecture, applied to personalized agents. The 2026 State of AI Agent Memory report (Mem0) and Best AI Agent Memory Frameworks in 2026 (Atlan) survey production-grade options: Mem0, Letta, Zep, Cognee, Memori. arxiv 2604.16548 Survey on the Security of Long-Term Memory in LLM Agents covers the security and integrity dimensions that matter for regulated DV flows.
  1. Episodic store: failure bundle archive. Every resolved bundle gets stored in a vector DB (Chroma, Qdrant, pgvector). Key fields: bundle hash, embedding, fingerprint, resolution file/line, resolving commit, owning team. Retrieved by semantic similarity when new bundles arrive.
  2. Semantic store: per-IP design book. A structured YAML/JSON per IP capturing invariants, conventions, protocol-specific rules. Updated on spec change. Loaded into the bundle composer's “known constraints” context block. Example fields for an AXI IP: burst_length_max, id_width, supported_burst_types, back_pressure_protocol.
  3. Procedural store: debug runbook. Per fingerprint class, the canonical first-checks. “For SCB_MISMATCH fingerprints with D3 in the state signature, first run +NO_D3_TRANSITIONS=1. If passes, the bug is power-gating related; consult pwr_controller.” The agent consults this before generating its own hypotheses.
  4. Memory consolidation cycle. Monthly: scan episodic memory for clusters (k-means on embeddings); promote recurring patterns into semantic memory. “47 past bugs involved CDC on clk_b” becomes “clk_b CDC is a known failure class — check first.” This is the maintenance most teams skip and most regret.
  5. Memory hygiene and audit. Stale episodic entries (bug fixed two release cycles ago, IP redesigned) should be pruned. Every memory write logs author + timestamp + provenance. For regulated flows, every memory consultation that influenced an agent decision is auditable after the fact.
Apply to DV
  • Start with episodic memory only. SQLite + a vector index over bundle hashes is a one-day build. Procedural and semantic stores can wait until you have enough episodic data to feed them.
  • The Full Pipeline post's Stage 6 (postmortem) is exactly the write path into episodic memory. If you have Stage 6, you already have episodic memory; the read path (similarity search in bundle composer) is the integration work.
  • Plan the memory consolidation cycle into your quarterly cadence. The episodic-to-semantic distillation is high-value but only if someone owns it — designate the owner before building the store.
  • For regulated DV flows, the audit dimension is not optional. Memory writes and consultations are first-class events in your structured log, with the same provenance discipline as the testbench logs themselves.

11. What Does Not Work (Honest Limits)

The anti-patterns. Each of these is a failure mode the research has measured and the practitioners have lived. Knowing them in advance is how you save the afternoon you would otherwise spend on a context strategy that was never going to work.

Research A composite synthesis of the failure findings cited across this post: lost-in-the-middle (2509.21361, Context Rot 2026), context bloat in long-horizon tasks (2601.07190 Active Context Compression), tool preference unreliability (2505.18135), cache-bypass anti-patterns (2601.06007 Don't Break the Cache), and memory security issues (2604.16548).
  1. Pasting raw regression logs into a chat. The textbook trigger of all three context-degradation mechanisms. The bundle pattern exists specifically to prevent this.
  2. Trusting mid-context information without verification. The agent “cited” a line that was 80K tokens deep in a 200K-token prompt. The lost-in-the-middle research says that citation is unreliable. Verify before trusting.
  3. Assuming long context equals good context. The marketing slide says 1M tokens; the effective window is closer to 130K. Cap your bundle composer at 50-60% of advertised, not 100%.
  4. Cache-bypass anti-patterns. Putting volatile content (timestamps, request IDs, dynamic counters) at the top of your prompt invalidates the cache on every request. Stable content first, volatile last — not negotiable.
  5. Unbounded memory growth. Episodic store grows forever; semantic store accumulates contradictory entries; procedural store fills with one-off recipes that never apply again. Plan for pruning from day one.
  6. Tool results dumped into context without summarization. The 50K-line log returned by run_smoke belongs in a file, not in the next turn's context. Always summarize tool outputs before the next reasoning step.
  7. One agent doing everything. The generalist agent over-fits to common cases and under-performs on rare ones. UVMarvel's multi-agent-per-protocol pattern is the prior art; specialize where the failure modes differ.
  8. Schema validation as the only check. Schema-valid is not semantically correct. A schema-valid hypothesis-rank response can still be three wrong hypotheses. Pair structured output with at least one of: smoke verification, human review, or formal proof.
Apply to DV
  • Print this list. The eight anti-patterns are the ones a junior engineer is most likely to learn the hard way; making them visible is cheaper than the lessons.
  • Add a context-engineering review item to your debug-pipeline retrospectives: which of these eight bit us this month? The recurring offenders tell you where to invest the next round of infrastructure.

Reading List

Context Engineering Foundations

  • Anthropic, Effective Context Engineering for AI Agents (September 2025) — the canonical industry essay
  • Anthropic, Writing Effective Tools for AI Agents (2025) — tool design principles
  • arxiv 2603.09619 — Context Engineering: From Prompts to Corporate Multi-Agent Architecture
  • LangChain blog, Context Engineering for Agents (2026)
  • State of Context Engineering 2026 (Aurimas Griciunas)

Effective Context Window & Degradation

  • arxiv 2509.21361 — Maximum Effective Context Window
  • arxiv 2407.03651 — Working Memory Test & Inference-time Correction
  • Morph LLM, Context Rot: Why LLMs Degrade as Context Grows (2026)
  • Stanford / UC Berkeley (2023), Lost in the Middle: How Language Models Use Long Contexts

Memory Systems

  • arxiv 2502.06975 — Episodic Memory is the Missing Piece for Long-Term LLM Agents
  • arxiv 2604.04853 — MemMachine: Ground-Truth-Preserving Memory
  • arxiv 2604.16548 — Survey on the Security of Long-Term Memory in LLM Agents
  • Mem0, State of AI Agent Memory 2026
  • Atlan, Best AI Agent Memory Frameworks in 2026

Context Compression

  • arxiv 2601.07190 — Active Context Compression (Focus architecture)
  • arxiv 2510.00615 — ACON: Optimizing Context Compression for Long-Horizon LLM Agents
  • arxiv 2510.08907 — Autoencoding-Free Context Compression via Semantic Anchors

Just-in-Time Retrieval & Tool Design

  • arxiv 2505.18135 — Tool Preferences in Agentic LLMs are Unreliable
  • arxiv 2412.04093 — Practical Considerations for Agentic LLM Systems
  • LogRocket, The LLM Context Problem in 2026

Prompt Caching

  • Anthropic prompt caching API documentation
  • OpenAI cached prefix pricing documentation
  • arxiv 2601.06007 — Don't Break the Cache: Prompt Caching for Long-Horizon Agentic Tasks

Structured Output

  • XGrammar (March 2026) — default backend for vLLM, SGLang, TensorRT-LLM
  • Outlines library — grammar-based generation with JSON Schema
  • OpenAI Strict Mode documentation (August 2024)
  • Anthropic native structured output documentation (late 2025)

Hardware Agent Context

  • arxiv 2512.06247 — DUET: Agentic Design Understanding via Experimentation
  • arxiv 2604.01572 — AI-Assisted Hardware Security Verification (IEEE VTS 2026)
  • Semiconductor Engineering, Human-Centered Agentic AI Comes to RTL Verification
  • EE Times, Agentic AI Tackles RTL Verification's Productivity Gap
  • Embedded.com, Siemens Agentic Toolkit Automates Chip Verification Workflows

Better Prompts for DV: A Researcher Guide to Prompt Engineering in Hardware Verification

Most prompting research is generic. Most DV blog content about AI is hype. This post sits in the small intersection: a researcher guide to prompt engineering techniques mapped explicitly to Design Verification problems, with research callouts naming the papers, numbered playbook patterns as the practitioner takeaway, and DV-application blocks grounding each technique in your UVM testbench or RTL workflow.

The differentiator is Section 10 — the systematic correlation between SWE prompting strategies (where the empirical research has been done) and the DV adaptations they imply. Section 12 closes with six ready-to-paste prompt templates, each annotated with the research pattern it borrows from.

A note on what this post is not. It is not a survey of LLM model capabilities (those change quarterly). It is not a tutorial on Anthropic vs OpenAI vs your local model (the workflow is provider-agnostic). It is not advocacy for AI replacing DV engineers (the research consistently shows the human-in-the-loop is what wins). It is the durable methodology layer beneath whichever model and provider you happen to use today.

1. Foundations: Zero-shot, Few-shot, Role

Before the fancy techniques, the basics that ground everything else. Each of the three primitives below has a clean role in DV; mixing them up is a leading source of frustration.

Research arxiv 2406.06608 The Prompt Report — PRISMA-grounded systematic survey of 58 prompting techniques. arxiv 2402.07927 Systematic Survey of Prompt Engineering. The 2025 MDPI review of 42 peer-reviewed SE studies clusters prompting into four research patterns: manual crafting, RAG, chain-of-thought, automated tuning. The Liu et al. 2026 taxonomy formalizes 58 distinct techniques across modalities.
  1. Zero-shot. A single instruction, no examples. Right for: tasks the model has clearly seen many times (generate a docstring, summarize a paragraph, classify by sentiment). Wrong for: anything that involves your team's idioms, your codebase's style, or a non-standard output schema.
  2. Few-shot in-context learning. 2-5 worked examples before the actual ask. The model picks up the shape of the desired output. Critically: examples should come from your own codebase, not from a generic library — the model already knows generic libraries.
  3. Role-prompting. "You are a senior DV engineer reviewing a 10-year-old testbench." The role routes the model into a more conservative, idiom-aware mode than the default helpful-assistant frame. The 2025 SE literature notes role prompts can be removed without much loss once you have good few-shot examples; until then, they are doing real work.
Apply to DV
  • Zero-shot for repetitive boilerplate: covergroup skeletons, factory registration, default field automation macros.
  • Few-shot for anything style-sensitive: your team's sequence library idioms, your scoreboard pattern, your monitor structure.
  • Role for tasks where the default helpful-assistant tone would be wrong: code review (assistant is too lenient), architectural critique (assistant defers), and triage (assistant hedges).

2. Chain-of-Thought for DV: The Caveat Section

CoT is the most famous prompting technique. It is also the one most likely to disappoint you on RTL tasks if you apply it naively. The fix is to scaffold the reasoning with HW-meaningful intermediate steps, not generic "let's think step by step."

Research Wei et al. 2022 (arxiv 2201.11903) — CoT activates reasoning in large models; without it even 540B-parameter models behave like much smaller ones. The 2026 RTL survey (Preprints 202509.1681) explicitly states that naive chain-of-thought has been largely ineffective in automating IC design workflows because the step granularity and reasoning direction do not align with expert RTL knowledge. arxiv 2504.06939 (FeedbackEval) shows that for code repair, removing structured-reasoning cues causes severe degradation.
  1. HW-aware CoT scaffolding. Replace "think step by step" with "walk through this cycle by cycle." The model needs the right intermediate representation; for hardware that is signal traces, FSM transitions, timing windows, and clock-domain crossings — not generic prose.
  2. Trace-conditioned CoT. If you have a waveform or structured log, paste the relevant signal slice into the prompt before asking for analysis. arxiv 2505.04441 shows trace-conditioned prompts consistently beat trace-free prompts for SWE code repair; the DV analog is direct.
  3. Bounded CoT. Ask for reasoning in N steps, not free-form. "Identify the failing cycle. Identify the violating signal. Identify the upstream cause. Propose the fix." Four bounded steps beat an open-ended "think about this."
  4. Avoid CoT for structural generation. RTL module generation rarely benefits from CoT — the model needs to emit a structured artifact, not reason about it. Use decomposition (Section 5) instead.
Apply to DV
  • HW-aware CoT for timing-related debug: assertion failures involving multi-cycle properties, arbitration races, CDC investigations.
  • Trace-conditioned CoT for any failure where the structured log captures the relevant signal evolution — combine with the JSON ring-buffer dump pattern.
  • Skip CoT entirely when asking for boilerplate or scaffolding; you want the artifact, not the model's commentary on producing it.

3. Self-Consistency and Self-Verification

A complex problem usually admits multiple correct reasoning paths. Sample many paths, vote on the consistent answer. The technique is dramatic on hard problems and almost free when the model supports temperature sampling.

Research Wang et al. 2022 (arxiv 2203.11171) introduces self-consistency: sample diverse reasoning paths, take the most consistent answer. arxiv 2510.01069 Typed Chain-of-Thought / Certified Self-Consistency (CSC, 2026) aggregates only over experiments satisfying typing constraints — 69.8% accuracy on GSM8K versus 19.6% baseline. arxiv 2603.08999 Confidence-Aware Self-Consistency maintains comparable accuracy with up to 80% fewer tokens. Self-Verification (Weng et al. 2022, arxiv 2212.09561): generate forward, verify backward.
  1. Vote on hypothesis-rank. When triaging a critical bug, run the hypothesis-rank prompt three times with temperature > 0. Take the hypothesis that appears highest in all three runs. Cheaper than it sounds; catches obviously-wrong single-shot answers.
  2. Verify forward, check backward. Generated a fix proposal? Ask the model to derive the test that would fail without the fix. If the model cannot, the proposal is suspect.
  3. Certified self-consistency for typed outputs. When the answer must conform to a schema (a coverage plan, a JSON fix proposal), aggregate only over candidates that pass type/schema validation. The 2026 CSC paper is the formal version of what good DV teams already do informally.
  4. Confidence-aware sampling. When the first two samples agree strongly, stop. When they disagree, sample more. The 2026 CASC paper formalizes the trade-off; the practical takeaway is to avoid blindly running N samples when 2-3 suffice.
Apply to DV
  • Self-consistency on RCA for silicon-escape bugs — the stakes justify the extra samples.
  • Self-verification on every proposed RTL fix before the human reads it: ask the model to derive a failing test for the unpatched version.
  • Certified self-consistency for coverage-plan generation: aggregate only over plans that parse against your covergroup schema.

4. Tree-of-Thoughts for Hard RCA

When chain-of-thought gets stuck on a wrong path, Tree-of-Thoughts explores multiple branches in parallel and prunes. The headline result is dramatic; the DV use case is hard root-cause analysis where 3+ hypotheses are equally plausible.

Research Yao et al. 2023 (arxiv 2305.10601) introduces Tree-of-Thoughts. The headline: on Game-of-24, GPT-4 with standard CoT solved 4% of problems; with ToT, 74%. arxiv 2401.14295 (Besta et al., Demystifying Chains, Trees, and Graphs of Thoughts) extends to graph-shaped reasoning. The pattern works because branching defers the commitment to a single reasoning trajectory.
  1. Branch the hypothesis space. "Generate three independent root-cause hypotheses, each starting from a different evidence anchor in the bundle." Force the model to start from different observations, not extend one chain of reasoning.
  2. Score and prune. After branching, ask the model to score each branch by "evidence weight" and propose which to prune. The model is often better at evaluating its own branches than picking one upfront.
  3. Iterate on the survivor. Take the highest-scored branch and apply CoT or self-debug only to that subtree. Saves the cost of exploring all branches deeply.
  4. Use when bug is "ambiguous." If a senior engineer cannot quickly name the most likely root cause from inspection, that ambiguity is the trigger for ToT. For obvious bugs, vanilla hypothesis-rank is cheaper.
Apply to DV
  • ToT on multi-IP SoC bugs where the failure could plausibly involve cache, fabric, or peripheral — let the model branch by subsystem.
  • ToT on protocol-error bugs where the violation could be the master, the slave, or the interconnect — branch by node.
  • Skip ToT for clear-cut bugs — the cost is real and the marginal benefit is zero when the answer is obvious.

5. Decomposition: Least-to-Most and Decomposed Prompting

Complex problems become a sequence of simpler ones. The technique transferred wholesale from theorem-proving to LLM prompting and remains one of the highest-leverage moves in your toolkit for any task larger than a single function.

Research Zhou et al. 2022 (arxiv 2205.10625) Least-to-Most Prompting — 16% to 99% on SCAN compositional generalization. arxiv 2210.02406 Khot et al. Decomposed Prompting — modular subproblems with composable solvers. The 2026 RTL research (PALM, arxiv 2506.09002) extends decomposition to program-analysis-derived path constraints; the DV analog uses coverage-bin-derived path constraints.
  1. Stage 1 — decompose. Ask the model to list the subproblems before solving any. "To verify this IP, list the 6-8 distinct sub-tasks in order of dependency." The decomposition itself is high-value output.
  2. Stage 2 — solve in order. Each subproblem solution becomes context for the next. The model maintains coherence because each step is bounded.
  3. Dependency-aware decomposition. When subtasks have non-linear dependencies, draw the DAG. The model can produce the DAG and then walk it in topological order.
  4. Path-constraint decomposition. For test generation: derive path constraints from coverage bins (analog of program-analysis-derived branching conditions in PALM), use each constraint as a sub-prompt.
Apply to DV
  • Verification plan: spec → coverage plan → sequence plan → scoreboard plan → checker plan → test list. Each step a separate prompt, each consuming the prior.
  • UVM agent scaffold: interface → transaction class → driver → monitor → sequencer → agent → example sequence. Each generated separately, each consistent with the prior.
  • SoC verification: top-level constraints → per-IP test plans → integration scenarios → stress tests. Top-down decomposition mirrors how senior engineers actually plan.

6. Few-Shot Done Right: Example Selection

Few-shot is the most common technique and the most commonly done badly. Random examples are sub-optimal. Text-similarity examples are sub-optimal. The 2024-2026 research has converged on better methods.

Research arxiv 2310.09748 LAIL (LLM-Aware ICL for Code Generation) — LLM labels candidate examples as positive (helpful) or negative (trivial); a model-aware retriever learns the preference. arxiv 2305.14210 Skill-Based Few-Shot Selection — pick examples that share the underlying skill needed, not just textual surface. arxiv 2412.02906 empirically: few-shot helps code synthesis but example quality dominates count.
  1. Examples from your codebase, not from the public web. The model already knows the public web. Your codebase carries your idioms, your naming, your error-handling conventions; the model picks these up from few-shot context cheaply.
  2. Structurally-similar examples beat textually-similar ones. Two AXI sequences that look different on the surface but share the same burst structure are better few-shot fodder than two cosmetically similar sequences with different structures.
  3. Negative examples are valuable. "Here is a similar task done wrong, here is the correct version." The contrast teaches the model what to avoid; bare positive examples cannot.
  4. 2-5 examples, no more. Beyond five, the marginal benefit per token drops sharply and the long-context degradation effect (arxiv 2509.21361) starts to bite.
Apply to DV
  • Maintain a curated examples directory in your TB repo: examples/sequences/, examples/scoreboards/, examples/checkers/. Each example is a known-good reference you few-shot from.
  • Index examples by protocol family and skill (AXI burst, AXI single, PCIe TLP, USB packet) so retrieval is structural.
  • Keep a small set of "canonical wrong" examples for negative few-shot — the seq that deadlocked, the scoreboard that missed an off-by-one. Contrast teaches.

7. Agentic Prompting: ReAct and Reflexion

When the LLM needs to iterate — observe, decide, act, observe again — you are in agentic territory. ReAct and Reflexion are the foundational patterns; the 2025-2026 HW research extends them with verifier-guided refinement.

Research Yao et al. 2022 (arxiv 2210.03629) ReAct — interleave Reasoning steps with Action calls (tool invocations) and Observations. Shinn et al. 2023 (arxiv 2303.11366) Reflexion — the agent self-critiques after each task and improves next time. arxiv 2509.06239 Proof2Silicon — verifier-guided prompt refinement via reinforcement learning, applied to verified hardware code generation.
  1. ReAct loop. Reason → Act (call a tool: query log, run smoke, fetch RTL excerpt) → Observe (tool output) → Reason. Three to five iterations beats one zero-shot answer almost universally for non-trivial problems.
  2. Bounded tool catalog. Per the 2025 Anthropic context-engineering essay: 4-6 self-contained tools beat one mega-tool. Tools must be self-contained, error-robust, and unambiguous. If a human cannot decide which tool to call, the agent cannot either.
  3. Reflexion after task completion. After the loop converges (or fails), ask the agent to write a paragraph about what worked, what did not, what to try differently next time. Cache the reflection; use it in the next session's system prompt.
  4. Verifier-guided refinement. Per Proof2Silicon: when a verifier (smoke test, formal property, lint rule) gives concrete feedback, route it back into the prompt as a hard signal. This is the HW-side analog of test-driven debug from DePro (arxiv 2603.19399): up to 64% fewer attempts and 7.6 minutes saved per problem in the SWE setting.
  5. Checkpoints and human approval. Per the 2025-2026 agentic-AI consensus: free-form agent loops are less reliable in production than graph-based orchestration with explicit state transitions, debuggability, and human-approval gates. Plan for the human checkpoint; do not assume autopilot.
Apply to DV
  • ReAct for triage: query_log, get_rtl_excerpt, run_smoke, query_fingerprint_db as the tool catalog. Anything else is scope creep.
  • Reflexion for postmortem: have the agent draft the "what we learned" section from its own trace. Edit for tone before publishing.
  • Verifier-guided refinement when you have a formal property: let the agent iterate on the property/fix until the formal tool stops complaining. The HW-side feedback signal is unusually strong.
  • Checkpoint on every fix proposal that touches RTL — never let the agent commit autonomously, regardless of how confident it sounds.

8. Constrained Decoding for Structured DV Output

When you need the LLM to produce a syntactically-valid artifact — JSON, Verilog, SVA, a coverage plan in your team's schema — you should be using constrained decoding rather than crossing your fingers. The infrastructure is cheap and the failure mode it eliminates is exactly the one DV engineers complain about most.

Research Constrained decoding modifies the LLM's sampling step via a logit processor: at each token position, valid tokens are computed from a grammar state and invalid tokens are masked. Libraries: llguidance (~50µs CPU per token; arbitrary context-free grammar), Outlines (compiles JSON schemas to O(1) lookup), Anthropic tool-call structured output, OpenAI JSON mode with response_format schemas. arxiv 2603.03305 Draft-Conditioned Constrained Decoding extends this to drafted generation with later refinement.
  1. JSON schema for structured outputs. Coverage plans, fix proposals, hypothesis-rank responses, test lists — all should land as schema-validated JSON, not as parseable-prose. The provider enforces the shape; you skip the defensive parser.
  2. Grammar-guided for HDL fragments. When asking for SVA or a small Verilog snippet, a grammar (or even a regex-shaped constraint) eliminates syntactic invalidity at the source. The model never emits the "almost valid" output that wastes a downstream tool invocation.
  3. Hybrid LLM + template. Per HAVEN (arxiv 2604.27643): the LLM produces a structured architectural plan in JSON; a rule-based generator emits the actual UVM. The split is what gets HAVEN to 100% compile success.
  4. Validate even with constrained decoding. Schema-validity is necessary, not sufficient. A schema-valid coverage plan can still be functionally wrong. Always pair constrained decoding with a downstream validation step.
Apply to DV
  • Define one JSON schema for each recurring DV output: coverage plan, fix proposal, hypothesis rank, test plan, refactor proposal. Reuse across teams.
  • For Verilog or SVA fragments, prefer hybrid generation: LLM produces a structured intermediate (signal list, property predicates, port map), template emits the syntactically-valid code.
  • Wire constrained-decoding errors into your alerting — if the model is failing constraint satisfaction repeatedly on a class of prompt, that is your signal the prompt itself needs work.

9. Meta-Prompting and Prompt Optimization

The prompts you write today are the next thing to refactor. Meta-prompting uses the LLM to rewrite your own prompts, A/B-tested against a small eval set. The 2026 industry consensus is that prompt engineering has matured from craft into versioned-and-tested engineering practice.

Research arxiv 2502.00728 Meta-Prompt Optimization for LLM-Based Sequential Decision Making. The 2026 industry literature (Comet, IntuitionLabs) reports self-refinement loops consistently improve prompt quality by 10-25%. arxiv 2503.02400 Promptware Engineering formalizes prompt engineering as a software engineering discipline: versioning, evaluation, regression testing for prompts.
  1. Build a small eval set. 15-25 representative DV prompts (debug, scaffold, review, refactor, cov-plan) with expected output shapes. This is your prompt regression suite.
  2. Self-refinement loop. Ask the LLM to critique and rewrite your prompt against the eval set. Compare A/B. Keep the winner. Repeat until the curve flattens.
  3. Version prompts like code. Prompts in a repo, semver tags, regression tests on every change. The day the underlying model updates, you re-run the suite to catch silent drift.
Apply to DV
  • Standardize the team's top 6-10 DV prompts in a shared repo. The hypothesis-rank prompt, the scaffolding prompt, the coverage-plan prompt — all versioned, all tested.
  • On every model upgrade or provider switch, re-run the eval suite. Catch the prompt that quietly stopped working before a junior engineer trusts a wrong answer.

10. SWE → DV: Methodological Correlations

This is the section that does not exist elsewhere on the internet. The SWE research community has done the empirical work on prompting for debug, test generation, code review, and refactoring. Each result transfers to DV with a clear adaptation. The table below names the SWE finding, the empirical evidence, the DV analog, and the specific adaptation needed.

Research Six recent SWE-prompting findings anchor this section. FeedbackEval (arxiv 2504.06939) on structured-reasoning code repair. DePro (arxiv 2603.19399) on iterative test-driven debug. PALM (arxiv 2506.09002) on program-analysis-derived path constraints. IntUT (paper covered in 2024-2025 SE literature) on intent-first unit-test generation hitting +94% branch coverage. Meta's semi-formal reasoning template hitting 93% code-review accuracy. Refactor subcategory explanation (arxiv 2411.02320) moving success from 15.6% to 86.7%. Static-analysis-augmented prompts (arxiv 2508.14419) cutting security violations from >40% to 13%.

The seven meta-patterns below each have an empirical SWE-side anchor and a direct DV adaptation. The DV adaptations are where the value is — they are not in the SWE papers because the SWE researchers were not thinking about hardware.

#SWE Meta-PatternSWE EvidenceDV Adaptation
1 Trace-Conditioned Prompting Execution traces in prompts consistently beat trace-free prompts (arxiv 2505.04441) Waveform-conditioned debug: paste FSDB signal slice formatted as "trace events"; Grove (arxiv 2511.x) confirms this for HW debug
2 Iteration-Feedback Loop DePro test-driven debug (64% fewer attempts); IntUT coverage-driven test gen (+94% branch coverage); FeedbackEval structured reasoning Smoke-driven self-debug loop; coverage-bin-driven test generation; formal-feedback loops (Proof2Silicon 2509.06239 is the HW-side prior art)
3 Intent-First Prompting IntUT: explicit test intentions improve branch coverage by 94%, line coverage by 49% Verification-intent-first: every UVM sequence prompt opens with "this is verifying scenario X / covering bin Y"; never "generate a test for this DUT"
4 Structured-Reasoning Templates Meta's semi-formal reasoning (premises → execution → conclusion) reached 93% code-review accuracy SVA-anchored review template: relevant assertions as premises; simulation/formal evidence as execution; coverage-bounded conclusion
5 Multi-Generation + Vote 5 refactorings per input → +28.8% pass rate; self-consistency over CoT samples Multi-fix-proposal + smoke gate: LLM proposes 3 fixes, smoke runs all 3, pick the one that passes. HW has a natural validator the SWE side often lacks — the simulator
6 Hybrid Tool+LLM Pipeline Static-analysis-augmented prompts (security >40% → 13%); PALM program-analysis path constraints Lint-augmented prompts; formal-augmented prompts; coverage-gap-augmented prompts. LASHED (arxiv 2504.21770) and Proof2Silicon are the HW-side prior art
7 Subcategory-Explanation Pattern Refactor success 15.6% → 86.7% from naming the refactoring type upfront (arxiv 2411.02320) UVM refactor categories upfront: extract base class, replace inheritance with composition, virtual-sequence extraction, agent split, factory-override simplification, config-object introduction. Name the category before asking for the change

Two threads tie the seven patterns together. First, every winning SWE prompting move has a feedback signal — tests, coverage, lint, formal — and the DV side has all of those signals available, often more cheaply than SWE does. Second, every winning move provides structure to the LLM upfront — the test intent, the refactor category, the schema — rather than asking the model to infer it. These two themes recur through Section 12's templates and Section 13's honest-limits discussion.

Apply to DV
  • Audit your top 10 prompts against the seven patterns. The patterns missing are usually the lowest-hanging adoption fruit.
  • For each adopted pattern, instrument a metric so you can measure the lift on your eval suite over time. The SWE results are large; the DV results should be too.

11. HW-LLM Frameworks Already in the Wild

The 2024-2026 hardware-LLM research has converged on a small number of named systems. Each one reveals a specific lesson about prompting for HDL/UVM that you can borrow without adopting the whole system.

Research Ten published HW-LLM systems, with the lesson each contributes: HAVEN (arxiv 2604.27643) plan-then-template, 100% compile, 90.6% coverage. UVM² (arxiv 2504.19959) domain-knowledge prompts + syntactic constraints + iterative refinement. UVMarvel (arxiv 2605.04704) multi-agent per protocol, 95.65% coverage. VeriGRAG (arxiv 2510.15914) structure-aware soft prompts. FVDebug (arxiv 2510.15906) causal graph synthesis + agentic exploration for waveform/RTL/spec debug. MEIC (arxiv 2405.06840) iterative debug with bounded progress per turn. LASHED (arxiv 2504.21770) LLM + static analysis for early RTL bug detection. Self-HWDebug (arxiv 2405.12347) self-instructing debug from vulnerable/secure RTL pairs. ReasoningV (arxiv 2504.14560) reasoning specialization for Verilog. Proof2Silicon (arxiv 2509.06239) RL prompt repair from formal feedback. CVDP (arxiv 2506.14074) benchmark establishes the 34% pass@1 ceiling.
  1. Never let the LLM emit HDL directly without scaffolding. This is the single lesson all 10 systems share. HAVEN's split (LLM → plan JSON; template engine → UVM) is the canonical pattern; the 100% compile success and 90.6% coverage numbers are the proof.
  2. Iterate with feedback, not in isolation. UVM², MEIC, and Proof2Silicon all use iterative refinement against a verifier signal (coverage, syntax, formal property). Single-shot generation is what produces the disappointing CVDP numbers; iteration is what closes the gap.
  3. Combine LLM with classical tools. LASHED pairs LLM with static analysis. Proof2Silicon pairs LLM with formal verification. The hybrid systems beat LLM-alone systems consistently.
  4. Specialize agents per protocol or subsystem. UVMarvel uses a different agent per bus protocol and reaches 95.65% coverage. Generalist agents over-fit to common protocols and under-perform on rarer ones.
  5. Borrow the patterns without adopting the systems. You do not need to deploy HAVEN to use its plan-then-template idea in your own prompts. The published systems are demonstrations; the patterns are reusable.
Apply to DV
  • Pick one HW-LLM lesson and apply it to your current workflow this month. Plan-then-template is the highest-leverage starter.
  • Track your own pass@1 number on a small local benchmark. The 34% CVDP ceiling is a public-research average; your number with structured prompts and iteration should be substantially higher on tasks scoped to your IP family.
  • When evaluating a vendor's AI-DV tool, ask which of the 10 patterns above it implements and which it skips. The honest vendors can answer.

12. A Concrete DV Prompt Template Library

Six ready-to-paste prompt templates, each annotated with the research pattern it borrows from. Replace the {{variables}} with your specifics. None of these templates are theoretical; each composes 2-3 of the techniques from Sections 1-10.

Template 1 — RTL Module from Spec (decomposition + few-shot + intent-first)

[SYSTEM]
You are a senior RTL designer writing SystemVerilog for a verification
team to validate. Match the style of the provided examples. Never invent
functionality not present in the spec.

[USER]
Generate a SystemVerilog module implementing the feature below.

Decomposition (handle in order):
1. Identify all signals required from the spec
2. Identify the FSM states (if any)
3. Identify the data flow (combinational vs sequential)
4. Generate the port list
5. Generate the module body
6. State any assumptions in a final comment

Spec section:
{{spec_excerpt}}

Module name: {{module_name}}
Style examples (from our codebase):
{{example_1}}
{{example_2}}

Why this works: The numbered decomposition (Section 5) prevents the model from emitting a monolithic blob. The codebase-derived examples (Section 6) carry team idioms the model would not infer. The explicit "never invent" instruction caps the hallucination surface.

Template 2 — UVM Agent Scaffolding (role + decomposition + few-shot + scope limits)

[SYSTEM]
You are a senior UVM verification engineer scaffolding a new agent.
Match the team's existing patterns. Never invent functionality not
explicitly asked for. Default to non-blocking, event-driven design.

[USER]
Generate UVM agent scaffolding for the interface below.

Interface definition:
{{interface_signature}}

Reference monitor (the team's style guide):
{{reference_monitor_code}}

Output, in order:
1. Transaction class (sequence_item) with rand fields and UUID stamping
2. Driver class with run_phase
3. Monitor class with run_phase and analysis_port
4. Sequencer typedef
5. Agent class with build_phase + connect_phase

Do NOT generate:
- The sequence library (separate prompt)
- The scoreboard (separate prompt)
- Any test class

Why this works: Role priming (Section 1) sets the assistant voice. Decomposition with explicit ordering (Section 5) keeps the agent components consistent. The "do NOT" list prevents the helpful-assistant scope creep that breaks generated TBs.

Template 3 — Failure RCA (Hypothesis-Rank) (CoT + multi-gen self-consistency + constrained-output JSON)

[SYSTEM]
You are a senior DV engineer triaging a UVM regression failure. Your
job is ranked hypotheses with falsifying experiments, not free-form
analysis. Reason from evidence in the bundle, never from generic
knowledge of the protocol.

[USER]
Analyze this failure bundle and propose the 3 most likely root causes
ranked by probability. For each cause:
- one-sentence explanation grounded in a specific bundle field
- ONE falsifying experiment (plusarg, variant, probe) runnable in <5 min
- the bundle field (line/event/signal) that supports the hypothesis

Return JSON only, no prose:
[
  {"cause": "...", "prob": 0.XX, "falsifier": "...", "evidence_field": "..."},
  ...
]

Bundle:
{{bundle_json}}

Why this works: Constrained-decoding output (Section 8) forces the structure to be machine-checkable. Forcing "grounded in a specific bundle field" closes the hallucination escape hatch. Run the same prompt with temperature > 0 three times for self-consistency (Section 3); accept the hypothesis appearing as #1 across all three runs.

Template 4 — Coverage Plan from Spec (decomposition + JSON schema + intent-first)

[SYSTEM]
You are a UVM verification engineer producing a coverage plan. Output
is JSON consumed by a covergroup generator. Cover features, not signals.
Every cross must have a reason.

[USER]
Generate a coverage plan from the spec section below.

Spec section:
{{spec_section}}

Decompose into:
1. Feature categories (top-level groups)
2. Per-feature bins (specific values to cover)
3. Per-feature crosses (combinations - each must have a reason)
4. Per-feature negative scenarios (error and corner cases)

Output JSON shape:
{
  "features": [
    {
      "name": "...",
      "bins": [{"name": "...", "values": [...]}],
      "crosses": [{"name": "...", "with": [...], "reason": "..."}],
      "negatives": [{"name": "...", "scenario": "..."}]
    }
  ]
}

Why this works: "Cover features, not signals" is intent-first prompting (Section 10 pattern 3) at the system level. The "every cross must have a reason" constraint prevents combinatorial blowup of meaningless crosses. JSON shape pinned for downstream tooling.

Template 5 — NL to SVA Translation (structured-reasoning template + bounded output)

[SYSTEM]
You are a formal verification engineer translating natural-language
requirements into SystemVerilog Assertions. Output must be syntactically
valid SVA. Use vacuous-pass-aware patterns. Use semi-formal reasoning:
state premises, trace execution, derive the property.

[USER]
Translate this requirement into one SVA property.

Requirement:
"{{nl_requirement}}"

Available signals:
{{signal_list_with_widths}}

Clock: clk (posedge), Reset: rst_n (active low)

Reasoning structure:
1. Premises: what signals participate and what their roles are
2. Execution: cycle-by-cycle trace of the requirement
3. Derivation: the SVA expression

Output:
- property_name: ...
- property body: ... (assert property)
- cover property: ... (proves the property is reachable, not vacuously true)
- non-vacuity note: one sentence explaining why this property cannot
  be vacuously true

Why this works: Meta's semi-formal reasoning template (Section 10 pattern 4) is the direct ancestor — premises, execution, derivation. The explicit non-vacuity check addresses the most common SVA bug: properties that pass because nothing triggers them.

Template 6 — UVM Refactor (Subcategory-First) (subcategory-explanation pattern + role + scope discipline)

[SYSTEM]
You are a senior UVM engineer refactoring legacy testbench code.
You must FIRST identify the refactoring category, THEN propose the
change. Never refactor without naming the category from the list below.

[USER]
Refactor this UVM code. The refactoring category is: {{category}}.

Categories available:
- extract_base_class: pull common functionality into a base class
- replace_inheritance_with_composition: convert is-a to has-a
- virtual_sequence_extraction: factor sequences from a monolithic test
- agent_split: separate driver/monitor/sequencer responsibilities
- factory_override_simplification: collapse N overrides into config_db
- config_object_introduction: replace string parameters with typed config

Code to refactor:
{{legacy_code}}

Output:
1. Confirm the category and explain in 2 sentences why it applies
2. The refactored code
3. Backwards-compatibility note: what tests or sequences need to change
4. Risk assessment: 1-10, with one sentence justification

Why this works: The subcategory-explanation pattern (Section 10 pattern 7) is what took refactoring success from 15.6% to 86.7% in the SWE literature. The explicit refactoring categories prevent the model from inventing its own. The risk-assessment line forces the model to consider what it is changing — not just produce the change.

Apply to DV
  • Adopt one template this week. The hypothesis-rank template (Template 3) has the highest leverage for engineers actively triaging regressions.
  • Version templates in your team repo. Add them to the eval suite from Section 9 so model upgrades cannot silently break them.
  • Build one new template per quarter as you discover prompts that consistently work for your team's workflow.

13. What Does Not Work in DV (Honest)

An honest accounting of where the research and the lived experience agree the LLM falls short. Knowing these in advance is what keeps you from wasting an afternoon on a prompt that was never going to work.

Research The 2026 RTL survey explicitly names naive CoT as ineffective for IC design. CVDP (arxiv 2506.14074) shows SOTA models hit only 34% pass@1 on hardware-specific tasks. ProtocolLLM (arxiv 2506.07945) shows even syntactic validity is unreliable on novel SystemVerilog testbench code. The effective-context-window paper (arxiv 2509.21361) shows degradation past ~1,000 tokens of operational context even for frontier models with advertised 200K windows.
  1. Naive chain-of-thought for RTL design. "Let's think step by step" produces generic reasoning that does not map to the actual structure of HDL synthesis. The fix is HW-aware CoT scaffolding (Section 2), not abandoning CoT entirely.
  2. Long raw log dumps as context. Pasting 2,000 lines of UVM log into a chat is the textbook trigger of the effective-window degradation. Use the JSON bundle pattern instead — 200 structured events beat 2,000 lines of prose every time.
  3. Asking for "novel" protocol implementations. The model knows AXI, PCIe, USB, and AHB because the training data does. Your team's proprietary or pre-release protocols are not in there; generated code looks right and is wrong.
  4. Trusting LLM-emitted line numbers. Frontier models routinely produce line numbers off by 4-10 lines, even for code in the prompt. Treat the line number as a neighborhood pointer, never as an exact location. The fix is to verify before applying.
  5. Random few-shot examples. The research is unambiguous: random or text-similarity examples are sub-optimal. If you do not invest in example selection (Section 6), you leave significant performance on the table.
  6. One-shot complex IP verification. "Verify this PCIe controller" as a single prompt produces vague boilerplate. Decomposition is non-negotiable for tasks larger than a single component.
  7. Believing published pass@1 numbers as your own ceiling. Benchmarks like CVDP report averages across heterogeneous tasks. Your number on tasks scoped to your IP family with structured prompts and iteration is almost certainly higher — measure it yourself.
  8. Single-sample answers for high-stakes decisions. Any output that goes into shipping silicon (an SVA the design relies on, a coverage point that gates signoff, an RTL fix) should pass either self-consistency, formal verification, or human review — not all three is fine, none of three is not.
Apply to DV
  • Print this section as a poster. The bottom three failure modes (LLM line numbers, one-shot IP verification, single-sample high-stakes) are the ones a junior engineer is most likely to learn the hard way.
  • For tasks the model genuinely cannot do (novel protocols, multi-cycle deep reasoning), stop trying. Fall back to traditional methods; do not waste your morning on iteration #4 of a prompt that was never going to converge.

Reading List

Prompting Foundations & Surveys

  • arxiv 2406.06608 — The Prompt Report (PRISMA-grounded survey of 58 techniques)
  • arxiv 2402.07927 — Systematic Survey of Prompt Engineering
  • arxiv 2407.12994 — Prompt Engineering Methods for NLP Tasks
  • Liu et al. 2026, Frontiers of CS — Comprehensive Taxonomy of Prompt Engineering Techniques

Reasoning (CoT, Self-Consistency, Tree-of-Thoughts)

  • arxiv 2201.11903 — Wei et al., Chain-of-Thought Prompting
  • arxiv 2203.11171 — Wang et al., Self-Consistency Improves CoT
  • arxiv 2305.10601 — Yao et al., Tree-of-Thoughts
  • arxiv 2401.14295 — Besta et al., Chains, Trees, Graphs of Thoughts
  • arxiv 2510.01069 — Typed CoT / Certified Self-Consistency (2026)
  • arxiv 2603.08999 — Confidence-Aware Self-Consistency (2026)

Decomposition & Few-Shot

  • arxiv 2205.10625 — Zhou et al., Least-to-Most Prompting
  • arxiv 2210.02406 — Khot et al., Decomposed Prompting
  • arxiv 2310.09748 — LAIL: LLM-Aware ICL for Code Generation
  • arxiv 2305.14210 — Skill-Based Few-Shot Selection
  • arxiv 2412.02906 — Does Few-Shot Help LLM Code Synthesis?

Agentic Patterns

  • arxiv 2210.03629 — Yao et al., ReAct
  • arxiv 2303.11366 — Shinn et al., Reflexion
  • arxiv 2509.06239 — Proof2Silicon: RL Prompt Repair from Formal Feedback

Constrained Decoding

  • arxiv 2603.03305 — Draft-Conditioned Constrained Decoding
  • llguidance (github.com/guidance-ai/llguidance) — ~50µs/token CFG enforcement
  • Outlines library — JSON-schema-compiled valid-token lookup

Meta-Prompting & Promptware Engineering

  • arxiv 2502.00728 — Meta-Prompt Optimization for Sequential Decision Making
  • arxiv 2503.02400 — Promptware Engineering

SWE Prompting Empirical Studies

  • arxiv 2504.06939 — FeedbackEval (code repair)
  • arxiv 2603.19399 — DePro (debug)
  • arxiv 2506.13186 — Empirical Evaluation of APR
  • arxiv 2505.04441 — Execution Traces for Program Repair
  • arxiv 2506.09002 — PALM (Rust unit test coverage)
  • arxiv 2407.00225 — Prompt Engineering for Unit Test Generation
  • arxiv 2402.00097 — Code-Aware Prompting for Coverage-Guided Tests
  • arxiv 2508.14419 — Static Analysis as Feedback Loop
  • arxiv 2411.02320 — Empirical Study on Code Refactoring (15.6% → 86.7%)
  • arxiv 2303.07839 — ChatGPT Prompt Patterns for Code Quality

HW-LLM Frameworks & Benchmarks

  • arxiv 2604.27643 — HAVEN (UVM TB synthesis, 100% compile, 90.6% coverage)
  • arxiv 2504.19959 — UVM² (LLM-aided UVM machine)
  • arxiv 2605.04704 — UVMarvel (subsystem-level, 95.65% coverage)
  • arxiv 2510.15914 — VeriGRAG (structure-aware soft prompts)
  • arxiv 2510.15906 — FVDebug (waveform + RTL + spec debug)
  • arxiv 2405.06840 — MEIC (iterative RTL debug)
  • arxiv 2504.21770 — LASHED (LLM + static analysis)
  • arxiv 2405.12347 — Self-HWDebug (self-instructing security debug)
  • arxiv 2504.14560 — ReasoningV (efficient Verilog generation)
  • arxiv 2506.14074 — CVDP benchmark (34% SOTA pass@1)
  • arxiv 2506.07945 — ProtocolLLM (SV testbench benchmark)
  • arxiv 2212.11140 — Benchmarking LLMs for Verilog (foundational)

Context Engineering

  • Anthropic, Effective Context Engineering for AI Agents (2025) — the canonical industry essay
  • arxiv 2509.21361 — Maximum Effective Context Window
  • arxiv 2603.04814 — Beyond the Context Window (fact-memory vs long-context)