Chip verification produces simulation traces that can overwhelm a language model's context window. The Back-to-the-Future paper tackles that access problem by converting waveform dumps into a relational database before agents inspect them. Je Yang, Ivan Lobov and Thomas Karpati propose BTTF as a way to answer engineering questions about signal behavior and connect anomalies with RTL source code.
Waveform access through SQLite
The offline preprocessing pipeline separates signal declarations from timestamped value changes. A signal-metadata table records identity and type; a signal-changes table stores transitions and references the corresponding declaration. Agents can then query scope, time and values without loading the entire waveform into their prompt.
The pipeline also offers pruning policies. Removing clock transitions reduces storage because those transitions dominate many dumps, while timestamps retain the timing information used by queries. Other policies discard signals with no transitions or remove signals whose activity falls below the design-wide average. Those options preserve different amounts of information, so the smallest database is not interchangeable with the most complete one.
In the reported preprocessing comparison, zero-toggle pruning reduces the footprint by 72.18% while retaining coverage over active nets. More aggressive below-average pruning reaches an 86.97% reduction, but the table reports only 64.23% coverage. Engineers would need to choose a policy consistent with the debugging question, rather than treat compression alone as success.
Agents share the query and audit work
An orchestrator divides the engineer's request among specialized agents. The Analysis Agent generates schema-aware SQL, while the Verification Agent ties anomalies to version-controlled RTL and checks assertion rules. An Evaluation Agent audits correctness and completeness, including whether queries follow the schema and use the intended timescale and scope.
The benchmark uses Gemini-2.5-Pro as the default model and a dataset containing 4.2 GB of VCD traces with 18,450 nets. The authors curate 150 natural-language questions with expert-annotated answers. BTTF reaches 95.33% overall execution accuracy in that experiment. Results vary by task: scope discovery scores 100%, while static assertion checks score 84.62%.
Accuracy carries a latency cost
The collaborative design receives a 17.3% higher score than a single-agent baseline on the authors' five-criterion evaluation. It also incurs an average 1.96x latency overhead. That is a trade-off for teams deciding whether the extra audit improves answers enough to justify slower responses.
This comparison differs from the Malena study of machine-learning harnesses, where added orchestration did not produce a statistically significant advantage over a capable coding-agent baseline. BTTF gives agents a specialized waveform database and verification roles; the two studies evaluate different workloads, so neither establishes a universal rule for agent complexity.
The paper remains a proof of concept on its annotated queries. The authors identify streaming ingestion, consistent scope-path normalization and a waveform-to-RTL symbol index as further work for larger designs. Its most useful contribution is a concrete data interface that allows engineers to test agent-assisted debugging against recorded signal behavior, with compression and response-time costs visible alongside accuracy.
Comments