Context Engineering for DV: From Failure Bundles to Agent Memory
In 2025, an estimated 65% of enterprise AI failures were attributed to context drift or memory loss during multi-step reasoning. That is not an edge case — it is the dominant failure mode. The field has responded by reorganizing around context engineering: the discipline of curating what tokens reach the model at each step, what the model remembers across steps, and how it retrieves what it needs without drowning in what it does not.
This post is the deep-dive on what that means for Design Verification. Eleven sections cover the effective-context-window research, the context taxonomy, minimum-viable-context discipline, memory systems for long-running DV agents, compression techniques, tool design, prompt caching economics, structured output as implicit context, the bundle anatomy that makes DV-specific context work, agent memory for long-running workflows, and the honest limits of all of it.
The structure follows the same pattern as the rest of this AI series — each section opens with a research callout citing the specific paper, numbered playbook patterns provide the takeaway, and DV-application blocks ground each technique in your UVM testbench or RTL workflow. Cross-links to the AI Playbook, DV Prompts, Full Pipeline, and Structured Logging posts thread throughout where appropriate.
- 1. The Effective Context Window Crisis
- 2. Context Taxonomy: What Lives in the Window
- 3. The Minimum-Viable-Context Principle
- 4. Memory Systems for DV Agents
- 5. Context Compression Techniques
- 6. Tool Design as Context Discipline
- 7. Prompt Caching Economics for DV Loops
- 8. Structured Output as Implicit Context
- 9. DV-Specific Context Patterns: Bundle Anatomy
- 10. Agent Memory for Long-Running DV Workflows
- 11. What Does Not Work (Honest Limits)
- Reading List
1. The Effective Context Window Crisis
The single most actionable finding the DV community has not internalized: advertised context windows are not effective context windows. The gap is large, the failure modes are predictable, and pretending otherwise is the #1 source of disappointing LLM-debug experiences.
- Measure your model's real MECW on your workload. Do not trust the marketing number. Build a 10-prompt eval set with varying context sizes (1K, 5K, 20K, 50K, 100K tokens) and measure accuracy degradation. The curve is what tells you where to cap.
- Place critical information at the start and end. The lost-in-the-middle research is unambiguous: middle-positioned facts are functionally ignored at large context sizes. The failing event in a bundle goes at the top; the system instructions stay at the top of the system prompt; long mid-prompt context dumps are anti-patterns.
- Cap at 50-60% of the advertised window. If your model claims 200K tokens, treat 120K as the soft ceiling and 60K as the comfortable working zone. The marginal token past 60% is doing less work than the first one.
- Avoid mid-context noise. Distractor interference — semantically similar but irrelevant content — is a measurable third degradation mechanism. The fix is not to include “helpful background” that is not load-bearing for the task.
- Re-measure on every model upgrade. MECW changes when the underlying model changes. The eval set from step 1 is your regression suite; run it on every provider update.
- Never paste raw regression logs into a chat. The 2,000-line log dump is the textbook trigger of all three degradation mechanisms.
- In your failure bundles (Section 9), the
failureevent lives at the top of the bundle JSON, before the 200-event context window. The agent sees the punchline first. - If your team uses a 200K model, set the bundle composer's
max_tokensat 60-80K, not 200K. The compose-and-pray approach kills accuracy long before it kills cost.
2. Context Taxonomy: What Lives in the Window
Not all context is the same. A clean taxonomy of what types of information compete for the window lets you optimize each independently — cache the stable, prune the ephemeral, retrieve the on-demand.
- System instructions. Stable across the session. Cache hot. This is your prompt template, role priming, output schema. Changes only on team-level prompt evolution.
- In-context examples (few-shot). Semi-stable. Swap the examples by task family (AXI debug, PCIe debug, scoreboard generation). Cache by task family; refresh when the team adds canonical examples.
- Retrieved knowledge. Ephemeral. Just-in-time. RTL excerpts, spec sections, past-bug context come in via tool calls when the agent asks for them — never preemptively loaded.
- Tool outputs. Ephemeral. Summarize aggressively and drop the raw form. The 800-line log slice from a
query_logcall should be reduced to 30 key events before the next reasoning step. - Conversation history. Semi-stable but compounds dangerously. Prune aggressively. At session checkpoints (every 8-10 turns), re-summarize the conversation into a 200-token state synopsis and let the rest fall off.
- Scratchpad / think tool. Ephemeral. Fresh per agent step. The reasoning output is not preserved across iterations — only the conclusions feed forward.
- System instructions: your senior-DV-engineer role priming, the bundle schema description, the output JSON shape. Cache hot.
- Examples: keep a per-protocol example library (axi, pcie, usb, custom). The bundle composer picks the relevant examples based on the failure's protocol family.
- Retrieved: RTL excerpts via
get_rtl_excerpt, past bugs viaquery_fingerprint_db, design constraints viaquery_spec. None loaded upfront. - Tool outputs:
run_smokereturns pass/fail + log digest, not the full log. The agent does not need 50K tokens of new log to learn the smoke passed.
3. The Minimum-Viable-Context Principle
The 2025-2026 industry consensus has shifted from “more tokens equal more capability” to just-in-time retrieval and progressive disclosure. The agent should fetch what it needs when it needs it, not start the session with everything that might be relevant.
- Lightweight references over data dumps. The bundle composer should emit
file_paths,signal_names,fingerprint_ids,commit_hashes— not the full contents. The agent fetches contents via tools when it decides the reference is relevant. - Tools for on-demand load.
get_rtl_excerpt(file, line_range),query_log(filter, max_events),get_commit_diff(hash). Each tool is the equivalent of grep/glob/read for the DV domain. - Progressive disclosure: broad to narrow. The first tool call returns a summary (e.g., 5 candidate RTL files); subsequent calls drill into the one the agent picks. Avoid round-trips on broad scans.
- Verify the agent did not over-ask. Cap tool-call budgets per session (8-12 calls maximum). When the agent is over budget, that is your signal the task scope is wrong — not that you should raise the budget.
- Your failure bundle (Section 9) carries identifiers, not data: failing transaction
txn_id, suspect signal names, candidate RTL file paths, fingerprint hash, run header. The agent calls tools to fetch what it actually needs. - For SoC-scale debug, the first tool call returns the per-IP failure summary; the agent picks the IP; subsequent calls drill into that IP only. Never pre-load all IPs.
- Cap your debug agent at 10 tool calls per session. If it needs more, the bundle was wrong or the task was too big — not the agent's fault, but yours.
4. Memory Systems for DV Agents
Memory has become a first-class architectural component — not an afterthought of the model's context window. The 2024-2026 research has settled on a four-type taxonomy borrowed from cognitive science, and the DV-side mapping is direct once you see it.
- Short-term memory (STM) is the context window itself. Do not fight it — treat it as scarce. Apply Sections 1-3 ruthlessly: cap context, prune mid-window noise, use just-in-time retrieval.
- Episodic memory: specific past experiences. The vector database of past failure bundles + their resolutions. When a new failure arrives, semantic-similarity-search the episodic store to find “we have seen something like this before.” The bundle composer surfaces matches as
fingerprint_siblings. - Semantic memory: distilled facts and constraints. Protocol invariants (AXI burst constraints, PCIe LTSSM rules), design constraints (clock-domain mappings, power-state transitions), team conventions (UVM agent patterns, naming standards). Updated rarely; consulted often.
- Procedural memory: learned strategies. Debug recipes that worked. “For SCB_MISMATCH fingerprints involving D3 power state, check power-controller reset domain first.” Built from the postmortem stage of the Full Pipeline; consulted by the bundle composer at agent invocation.
- Memory consolidation: episodic to semantic over time. The MemMachine pattern: episodic stores raw experience; a periodic consolidation pass distills recurring patterns into semantic facts. For DV: monthly job that converts “these 47 past bugs all involved CDC issues on clk_b” into a semantic-memory entry “clk_b CDC is a frequent failure class — check first.”
- Stand up an episodic memory store (SQLite + vector index): one row per resolved fingerprint, embedding of the bundle, link to the fix commit. The Full Pipeline post's Stage 6 already writes this; this is its purpose.
- Stand up a semantic memory store: a structured YAML or JSON file per IP capturing “known invariants” (signal width assumptions, FSM transition rules, power-state semantics). Updated on spec change; consulted on every debug session.
- Stand up a procedural memory store: a runbook indexed by failure-class fingerprint — “for fingerprint class X, run plusarg Y first.” Build this empirically from successful debug sessions.
- Memory consolidation: a quarterly review that asks “what episodic entries clustered into recurring patterns?” and writes them into semantic memory. This is the maintenance step most teams skip and most regret.
5. Context Compression Techniques
Compression is the engineering response to the effective-window crisis: keep the load-bearing tokens, shed the rest. The 2026 research has converged on three concrete approaches, each with measured token-reduction numbers in the 26-54% range without task-success degradation.
- Anchored iterative summarization. At checkpoints, summarize prior turns into a 200-token state synopsis anchored to key entities (failure event, hypothesis, last tool result). Subsequent turns reference the synopsis, not the raw history.
- Failure-driven guideline optimization (ACON). When the agent fails to converge, the meta-loop asks “what part of the context was load-bearing? What was noise?” and updates the compression policy for next time. Compression learns from misuse.
- Autonomous pruning (Focus). Agents decide for themselves when to consolidate raw history into a persistent “knowledge block” and drop the rest. The slime-mold metaphor: keep the trails to food, prune the dead ends.
- Semantic anchor compression (SAC). Instead of summarizing prose, compress around named entities (the failing signal, the suspect FSM, the relevant txn_id) and let the model reconstruct details from anchors at inference time.
- Provider-native compaction. Anthropic and OpenAI both offer compaction APIs that summarize prior turns automatically. Cheaper than DIY; less control. Use when your eval suite says quality is preserved.
- For long debug sessions: at every 5 tool calls, re-summarize the session into a state synopsis — failing event, current hypothesis, what has been falsified, what is unknown. The next 5 calls operate on the synopsis, not the raw trace.
- Compress regression history into per-fingerprint summaries rather than carrying every past failure. A summary line per fingerprint with hit-count and last-resolution outperforms the raw event stream for context purposes.
- Use semantic anchors for DV: the
txn_idis the natural anchor for a transaction; the fingerprint is the natural anchor for a failure class; the signal name is the natural anchor for a debug investigation. Compress around these.
6. Tool Design as Context Discipline
Tool design is context engineering by other means. A well-designed tool returns minimum-viable context; a poorly-designed tool floods the window with noise and burns the budget you spent the previous five sections protecting. Anthropic's 2025-2026 guidance is unusually concrete.
- High-leverage tools, not thin wrappers. A tool that expands agent capability is worth a tool slot; a tool that just relays an API call is not.
query_log(filter, max_events)with a focused filter beatsread_file(path)applied to a 50MB log. - Clear, distinct names. No two tools should be plausibly confusable.
get_rtl_excerptvsread_rtl_fileis exactly the confusion to avoid — pick one verb, document the contract, deprecate the other. - Human-readable fields beat raw IDs. Return
{"signal": "axi_wb_inflight", "value": 0, "at_ns": 4960}, not{"s": 42, "v": 0, "t": 4960}. The agent reasons about names, not opaque integers. - Pagination, truncation, filtering as defaults. Every tool should return at most a few hundred tokens by default. A tool that could return 50K tokens must have explicit pagination parameters and the agent must opt in.
- Defensive against misuse. Validate parameters at the tool boundary; return clear error messages the model can recover from. The agent will pass nonsense; your tool should reject it gracefully without filling the window with stack traces.
- The DV agent tool catalog:
query_log(filter, max_events=200),get_rtl_excerpt(file, line, context_lines=20),run_smoke(test, seed, plusargs, timeout=300),query_fingerprint_db(fingerprint),get_commit_diff(hash),list_signal_history(signal, time_range). Six tools, each with explicit limits. - Tool outputs are digests, not raw artifacts.
run_smokereturns{pass: true/false, duration_s, error_class, last_10_events}, not the full 50K-line simulation log. - If a tool has more than 5 parameters, split it. If a tool's name needs a paragraph to explain, rename it.
7. Prompt Caching Economics for DV Loops
Prompt caching is the highest-ROI cost lever for any agentic workload that re-queries the same context dozens of times per session. UVM debug loops are exactly that workload. For teams running this at scale, not using prompt caching is not a cost optimization to skip — it is unaffordable not to use.
- Order context for cacheability. Stable content first, volatile content last. The system prompt, role priming, output schema, and codebase examples go at the top — they do not change between requests in the session. The failure event and per-request bundle slice go at the bottom — they change every call.
- Cache breakpoint design (Anthropic). Anthropic gives you explicit cache breakpoints with configurable TTL (5 minutes default). Place breakpoints after each stable block. The agent gets a 10% read on everything before the breakpoint, full price only on the new tokens after.
- Cache invalidation strategy. When does the cache expire? When the system prompt changes, when the schema version bumps, when the team examples are refreshed. Plan invalidation around team release cadence, not silently.
- Track cache hit rate as a first-class metric. If your agent's cache hit rate is below 60%, your context ordering is wrong. If it is above 90%, you are doing context engineering right.
- For UVM debug sessions where the agent re-queries the same bundle context across 5-15 ReAct iterations, prompt caching cuts per-iteration cost by ~90% on cached tokens. A $0.30 debug session becomes a $0.05 debug session at no quality cost.
- Order your debug agent context:
[stable: system prompt + role + bundle schema + protocol examples]— cache breakpoint —[volatile: failure event + 200 context events + tool outputs from this turn]. - Add cache-hit-rate to your debug-pipeline dashboard. If you do not measure it, the cache is silently degrading and you will discover only when the bill arrives.
8. Structured Output as Implicit Context
Schema-enforced structured output is not just a parsing convenience — it is implicit context. The schema you pin to the response tells the model exactly what fields to produce, in what order, with what types. That schema is doing context-engineering work without consuming context tokens.
- Three reliability levels — know which one you are using. If your prompt says “return JSON with fields X, Y, Z” you are at level 1 (80-95%). If you use the provider's tool-call interface, you are at level 2 (95-99%). If you use native structured output with FSM enforcement, you are at level 3 (100% schema-valid).
- Schema-as-context. The schema describes the output shape and the model treats it as instruction. A well-named schema field (
fingerprint_hash) is more useful than an instruction-prose equivalent (“include a fingerprint hash here”). - Schema versioning is part of context engineering. When you change a schema, you change the implicit context. Version your schemas; include the version in cache invalidation; include it in your prompt eval suite.
- Validate even with constrained decoding. Schema-valid does not mean semantically correct. A schema-valid coverage plan can still be a bad coverage plan. Constrained decoding eliminates one class of errors; downstream validation handles the rest.
- Maintain a DV schema library:
hypothesis_rank.schema.json,coverage_plan.schema.json,fix_proposal.schema.json,refactor_proposal.schema.json. Version each. - Use level-3 native structured output for any agent response that feeds downstream tooling. Use level-1 only when humans are the only consumer.
- The schema definitions themselves are part of your team's context-engineering surface area. Treat them with the same review discipline as production code.
9. DV-Specific Context Patterns: Bundle Anatomy
Everything above sets up this section. The DV-specific context pattern — the failure bundle — is what makes context engineering land for hardware verification. This is the deep-dive on bundle anatomy, RTL excerpt selection, signal slice formatting, DUT state representation, and the multi-IP coordination patterns that emerge at SoC scale.
- Bundle anatomy — six well-typed sections. The DV failure bundle is a JSON document with: (a)
failure— the triggering event, (b)context_events— ring-buffer dump of pre-failure events (Structured Logging Pattern 9), (c)rtl_excerpts— relevant RTL code snippets with reasons for inclusion, (d)fingerprint_siblings— past bugs with the same fingerprint plus their resolutions, (e)recent_commits— RTL git log for the last N days scoped to the failing IP, (f)run_header— seed, plusargs, tool version, RTL hash for replay. - RTL excerpt selection — symbol-indexed, not text-grep. Build a Verible / sv-parser / Surelog index of your RTL once, cached. For each failure, the bundle composer queries the index by signal names appearing in the context events, scoped to the failing IP. Returns the file path, line range, and 25 surrounding lines for each match, ranked by hit count. Never blind grep across the SoC.
- Signal slices as DUT_STATE events. Your
bind-based harness (see Debug page Pattern 11) emitsDUT_STATEJSONL events on key triggers: FSM transitions, queue depth changes, register writes. These appear as first-class events in the context_events stream — the agent sees them inline, not as a separate waveform query. - DUT state representation. A snapshot DUT_STATE event contains the current state of the things that change rarely but matter critically: FSM states, queue depths, key configuration registers, power state, clock domain, mode/profile. Format as flat key-value pairs in the JSONL event; the agent reads them with zero ceremony.
- Per-IP bundle scoping for SoC debug. At SoC scale, never compose a bundle that spans IPs by default. The first-pass bundle is per-IP; cross-IP correlation happens via shared
txn_idthreading (Structured Logging Pattern 3). The agent decides when it needs to cross IP boundaries and fetches additional per-IP bundles via tool calls. - MCP-based context exposure (Siemens pattern). The Model Context Protocol gives you a standardized server interface for exposing DV tools to any MCP-compatible client (Claude Code, IDE assistants, custom agents). Each tool from Section 6 becomes an MCP endpoint; the protocol handles the rest.
- Bundle hash for deterministic replay. Every bundle gets a content hash. Cache prompt+response by hash. The same failure replays to the same agent response, every time. Reproducibility is non-negotiable when the agent's output goes anywhere near production silicon decisions.
- If you build only one piece of context-engineering infrastructure this quarter, build the bundle composer. Everything else — agent design, memory, prompt caching — assumes the bundle exists.
- Adopt MCP early. The 2026 vendor landscape (Siemens, Cadence experimental, Synopsys announced) is converging on MCP as the integration protocol. Build your DV tool catalog as MCP-compatible from day one.
- The bundle composer is the single most important piece of Python in your DV-AI stack. Treat it accordingly: code review, unit tests, schema versioning, performance budget.
- For multi-team SoC environments, agree on a shared bundle schema across teams. The schema is a contract; each IP team produces bundles to it, agents consume bundles in the same shape regardless of source.
10. Agent Memory for Long-Running DV Workflows
Section 4 laid out the memory taxonomy. This section grounds it in the long-running workflows DV teams actually run: regression seasons that span weeks, projects that span quarters, IP families that span years. Memory is what turns the agent from a smart one-shot tool into a teammate that gets better with the project.
- Episodic store: failure bundle archive. Every resolved bundle gets stored in a vector DB (Chroma, Qdrant, pgvector). Key fields: bundle hash, embedding, fingerprint, resolution file/line, resolving commit, owning team. Retrieved by semantic similarity when new bundles arrive.
- Semantic store: per-IP design book. A structured YAML/JSON per IP capturing invariants, conventions, protocol-specific rules. Updated on spec change. Loaded into the bundle composer's “known constraints” context block. Example fields for an AXI IP:
burst_length_max,id_width,supported_burst_types,back_pressure_protocol. - Procedural store: debug runbook. Per fingerprint class, the canonical first-checks. “For SCB_MISMATCH fingerprints with D3 in the state signature, first run
+NO_D3_TRANSITIONS=1. If passes, the bug is power-gating related; consult pwr_controller.” The agent consults this before generating its own hypotheses. - Memory consolidation cycle. Monthly: scan episodic memory for clusters (k-means on embeddings); promote recurring patterns into semantic memory. “47 past bugs involved CDC on clk_b” becomes “clk_b CDC is a known failure class — check first.” This is the maintenance most teams skip and most regret.
- Memory hygiene and audit. Stale episodic entries (bug fixed two release cycles ago, IP redesigned) should be pruned. Every memory write logs author + timestamp + provenance. For regulated flows, every memory consultation that influenced an agent decision is auditable after the fact.
- Start with episodic memory only. SQLite + a vector index over bundle hashes is a one-day build. Procedural and semantic stores can wait until you have enough episodic data to feed them.
- The Full Pipeline post's Stage 6 (postmortem) is exactly the write path into episodic memory. If you have Stage 6, you already have episodic memory; the read path (similarity search in bundle composer) is the integration work.
- Plan the memory consolidation cycle into your quarterly cadence. The episodic-to-semantic distillation is high-value but only if someone owns it — designate the owner before building the store.
- For regulated DV flows, the audit dimension is not optional. Memory writes and consultations are first-class events in your structured log, with the same provenance discipline as the testbench logs themselves.
11. What Does Not Work (Honest Limits)
The anti-patterns. Each of these is a failure mode the research has measured and the practitioners have lived. Knowing them in advance is how you save the afternoon you would otherwise spend on a context strategy that was never going to work.
- Pasting raw regression logs into a chat. The textbook trigger of all three context-degradation mechanisms. The bundle pattern exists specifically to prevent this.
- Trusting mid-context information without verification. The agent “cited” a line that was 80K tokens deep in a 200K-token prompt. The lost-in-the-middle research says that citation is unreliable. Verify before trusting.
- Assuming long context equals good context. The marketing slide says 1M tokens; the effective window is closer to 130K. Cap your bundle composer at 50-60% of advertised, not 100%.
- Cache-bypass anti-patterns. Putting volatile content (timestamps, request IDs, dynamic counters) at the top of your prompt invalidates the cache on every request. Stable content first, volatile last — not negotiable.
- Unbounded memory growth. Episodic store grows forever; semantic store accumulates contradictory entries; procedural store fills with one-off recipes that never apply again. Plan for pruning from day one.
- Tool results dumped into context without summarization. The 50K-line log returned by
run_smokebelongs in a file, not in the next turn's context. Always summarize tool outputs before the next reasoning step. - One agent doing everything. The generalist agent over-fits to common cases and under-performs on rare ones. UVMarvel's multi-agent-per-protocol pattern is the prior art; specialize where the failure modes differ.
- Schema validation as the only check. Schema-valid is not semantically correct. A schema-valid hypothesis-rank response can still be three wrong hypotheses. Pair structured output with at least one of: smoke verification, human review, or formal proof.
- Print this list. The eight anti-patterns are the ones a junior engineer is most likely to learn the hard way; making them visible is cheaper than the lessons.
- Add a context-engineering review item to your debug-pipeline retrospectives: which of these eight bit us this month? The recurring offenders tell you where to invest the next round of infrastructure.
Reading List
Context Engineering Foundations
- Anthropic, Effective Context Engineering for AI Agents (September 2025) — the canonical industry essay
- Anthropic, Writing Effective Tools for AI Agents (2025) — tool design principles
- arxiv 2603.09619 — Context Engineering: From Prompts to Corporate Multi-Agent Architecture
- LangChain blog, Context Engineering for Agents (2026)
- State of Context Engineering 2026 (Aurimas Griciunas)
Effective Context Window & Degradation
- arxiv 2509.21361 — Maximum Effective Context Window
- arxiv 2407.03651 — Working Memory Test & Inference-time Correction
- Morph LLM, Context Rot: Why LLMs Degrade as Context Grows (2026)
- Stanford / UC Berkeley (2023), Lost in the Middle: How Language Models Use Long Contexts
Memory Systems
- arxiv 2502.06975 — Episodic Memory is the Missing Piece for Long-Term LLM Agents
- arxiv 2604.04853 — MemMachine: Ground-Truth-Preserving Memory
- arxiv 2604.16548 — Survey on the Security of Long-Term Memory in LLM Agents
- Mem0, State of AI Agent Memory 2026
- Atlan, Best AI Agent Memory Frameworks in 2026
Context Compression
- arxiv 2601.07190 — Active Context Compression (Focus architecture)
- arxiv 2510.00615 — ACON: Optimizing Context Compression for Long-Horizon LLM Agents
- arxiv 2510.08907 — Autoencoding-Free Context Compression via Semantic Anchors
Just-in-Time Retrieval & Tool Design
- arxiv 2505.18135 — Tool Preferences in Agentic LLMs are Unreliable
- arxiv 2412.04093 — Practical Considerations for Agentic LLM Systems
- LogRocket, The LLM Context Problem in 2026
Prompt Caching
- Anthropic prompt caching API documentation
- OpenAI cached prefix pricing documentation
- arxiv 2601.06007 — Don't Break the Cache: Prompt Caching for Long-Horizon Agentic Tasks
Structured Output
- XGrammar (March 2026) — default backend for vLLM, SGLang, TensorRT-LLM
- Outlines library — grammar-based generation with JSON Schema
- OpenAI Strict Mode documentation (August 2024)
- Anthropic native structured output documentation (late 2025)
Hardware Agent Context
- arxiv 2512.06247 — DUET: Agentic Design Understanding via Experimentation
- arxiv 2604.01572 — AI-Assisted Hardware Security Verification (IEEE VTS 2026)
- Semiconductor Engineering, Human-Centered Agentic AI Comes to RTL Verification
- EE Times, Agentic AI Tackles RTL Verification's Productivity Gap
- Embedded.com, Siemens Agentic Toolkit Automates Chip Verification Workflows