The AI-for-DV Research Reading List (2024-2026)

This is the citation home for the AI page. Every research claim in the Practitioner Playbook and its deep-dive posts traces back to an entry here. The list is grouped by theme, spans 2024–2026, and gets updated as the field moves.

Two kinds of entries: papers you should actually read (marked in the Start Here box), and papers you should know exist so you can find them when a problem lands on your desk. Where a named system has no standalone paper link, the entry points at the survey that covers it.

Start Here — Three Papers

  • The ASPDAC 2026 survey (Surveys, below) — the best current map of LLM-assisted verification: assertion generation, testbench automation, and RTL debug in one method tree.
  • CVDP (Code Generation, below) — the 783-problem hardware benchmark behind the 34% pass@1 ceiling. Read it before believing any codegen demo.
  • The Prompt Report (Prompt & Context, below) — the systematic survey the practical prompting patterns are drawn from.

Surveys & Field Maps

  • LLM-Assisted Circuit Verification: A Comprehensive Survey (ASPDAC 2026) — the field map: SVA generation (prompting / RAG / training-based branches), testbench and test automation, automated RTL debugging, and collaborative verification frameworks.
  • arxiv 2512.23189 — The Dawn of Agentic EDA — three-tier method taxonomy (prompt-based, fine-tuned, multi-agent), the "unit-test fallacy" critique of module-level benchmarks, and the call for an Open Agentic EDA Standard.

Prompt & Context Engineering

  • Anthropic, Effective Context Engineering for AI Agents (2025) — the canonical industry essay on the shift.
  • arxiv 2406.06608 — The Prompt Report, 58 techniques, PRISMA-grounded.
  • arxiv 2402.07927 — Systematic Survey of Prompt Engineering.
  • arxiv 2407.12994 — Prompt Engineering Methods for NLP Tasks survey.
  • arxiv 2509.21361 — Maximum Effective Context Window.
  • arxiv 2603.04814 — Beyond the Context Window, fact-based memory vs. long-context.
  • arxiv 2601.01954 — Reporting LLM Prompting in Automated SE.

Pair Programming & Debugging

Code Generation

  • arxiv 2604.24621 — Evaluation of LLM-Based SE Tools.
  • arxiv 2503.01245 — LLMs for Code Generation Comprehensive Survey.
  • arxiv 2505.02133 — Multi-Agent Collaboration and Runtime Debugging.
  • arxiv 2506.14074 — CVDP, 783-problem hardware benchmark.
  • arxiv 2506.07945 — ProtocolLLM, SV testbench benchmark.

Assertions & Formal Verification

  • arxiv 2410.23299 — FVEval (NVIDIA) — the first comprehensive benchmark for LLMs on formal verification: NL-to-SVA translation and direct assertion suggestion from RTL, in graded tiers.
  • AssertLLM (ASPDAC 2025) — multi-LLM pipeline generating SVAs directly from full specification documents; reported 89% of generated assertions syntactically and functionally correct on its evaluated design.
  • AutoSVA2, ChIRAAG, LASSO, AssertionForge, Hybrid-NL2SVA — the prompting / RAG / fine-tuned branches of the SVA-generation method tree; see the ASPDAC 2026 survey above for the full map.

Coverage & Benchmarks

  • Revisiting VerilogEval (ACM TODAES) — a year of LLM progress on the canonical RTL-generation benchmark, extended to spec-to-RTL tasks and failure classification.
  • arxiv 2311.00176 — ChipNeMo (NVIDIA) — domain-adapted LLMs for chip design; the reference point for the fine-tuned-model route.
  • LLM4DV (FCCM 2025) and VerilogReader (LAD 2024) — coverage-directed stimulus generation with feedback loops on uncovered bins; both covered in the ASPDAC 2026 survey above.

Technical Debt

  • arxiv 2601.06266 — SATD in LLM Software, three new debt types.
  • arxiv 2507.03536 — ACE: Validated LLM Refactorings.
  • arxiv 2501.09888 — Automated SATD Repayment.

Agentic Design Patterns

  • arxiv 2601.12560 — Agentic AI Architectures & Evaluation.
  • arxiv 2510.09244 — Fundamentals of Building Autonomous LLM Agents.
  • arxiv 2604.00835 — Agentic Tool Use.
  • arxiv 2510.25445 — Agentic AI Survey.
  • arxiv 2604.27643 — HAVEN, UVM testbench synthesis.
  • arxiv 2504.19959 — UVM².
  • arxiv 2605.04704 — UVMarvel, subsystem-level UVM testbench construction.

Security & IP Safety

  • arxiv 2604.01572 — VTS survey of AI-assisted hardware security verification — the five-stage pipeline from asset identification through countermeasure reasoning; covers SV-LLM and SoCureLLM (HOST 2025).
  • arxiv 2503.13116 — IP leakage from fine-tuning on in-house Verilog; reports up to 46.52% of generated code similar to the source IP.
  • arxiv 2405.07061 — LLMs and the Future of Chip Design — security risks and building trust in AI silicon flows.

Limits & Pitfalls

  • arxiv 2411.09916 — "Should I Give Up Now?" LLM Pitfalls in SE.

Conference Proceedings Worth Browsing

  • DVCon US 2025 proceedings — agentic verification, coverage closure with AI, and formal + GenAI flows from practitioners.
  • DVCon Europe 2025 program — three full sessions on AI in verification, including an industry-practice paper on RL-driven coverage closure.
Author
Milan Kubavat
Sharing knowledge about silicon verification, hardware design, and engineering insights.

Comments (0)

Leave a Comment