CASE STUDY · AI-ASSISTED RTL ENGINEERING

From task to a board-tested LTE Turbo decoder in two days

In about 33 hours, an AI-assisted workflow produced a synthesizable LTE Turbo decoder, changed the hardware architecture, closed implementation, and completed a controlled board check. This is a record of one bounded engineering case, including the result that did not improve.

AlgoSilicon engineering · September 2026 · 8 min

The useful result is a reviewable hardware candidate with an implementation report, an independent expected stream, a multi-billion-cycle board record, and a deliberately broken build that the checker rejected.

What “two days” includes

The clock started with the task definition on 23 September. Board playback and its refutation control closed 28 hours, 33 minutes later; the current evidence package closed after 33 hours, 4 minutes. In that interval the workflow had to understand the LTE Turbo structure, complete an implementation after the initial generation pass did not finish the SISO datapath, choose a resource-aware architecture, generate RTL, run independent comparisons, close post-route timing and put the stream through hardware.

33 h 04 min
Task start to the current evidence package for this case MEASURED

This is a case-specific elapsed time, not a promise that every communications IP takes two days. The repeatable asset is the evidence-producing workflow: a candidate does not graduate because it looks plausible.

The first pass did not complete the decoder

Turbo decoding is a poor test for surface-level code generation. The constituent trellis is straightforward, but a useful implementation must coordinate forward and backward metrics, systematic and parity observations, QPP interleaving, extrinsic feedback, frame boundaries and iterative scheduling. The initial generation pass stopped short of a working SISO/Turbo implementation. The architecture stage then completed the missing datapath and scheduling.

That distinction matters. The process did not hide a failed attempt or simply keep prompting until a simulation turned green. It changed the level of reasoning from syntax to hardware ownership: what is stored, what is reused, what identifies the current frame and half-iteration, and which state is permitted to advance.

The architecture decision was time reuse

A naive six-iteration structure can imply twelve copies of the half-iteration machinery. The candidate instead reuses one Max-Log-MAP SISO engine across twelve half-iterations. Frame memories supply natural or QPP order, while extrinsic data returns with enough identity to avoid mixing phases or frames.

Simplified LTE Turbo decoder architecture decision
The core architectural choice: time-reuse the expensive SISO engine and make ordering and feedback explicit. Conceptual view, not an RTL schematic.

The current implementation accepts all 188 LTE QPP sizes from K=40 to K=6144. The measured flagship point uses K=6144, six iterations, WINDOW=32 and 5-bit channel inputs. It is a decoder core, not a claim for CRC, rate matching, HARQ, segmentation or a complete LTE modem.

Optimization means keeping the cost you did not win

On xczu67dr-fsve1156-2-i with the same out-of-context route flow, the candidate closed at 450.045 MHz using 2,657 LUTs, 1,177 flip-flops, 6.5 BRAMs and no DSPs. The MathWorks R2026a baseline closed at 396.354 MHz using 3,446 LUTs, 4,677 flip-flops, 7.0 BRAMs and no DSPs.

The candidate therefore uses less logic and memory and closes at a higher clock. It also reaches the first decision 684 cycles later: 75,379 cycles against 74,695. That loss stays in the record. The measured candidate frame throughput is 31.539 Mb/s; the baseline figure is only an upper bound of 28.0 Mb/s because its next-frame acceptance boundary was not recorded. The honest conclusion is that neither point dominates.

Same-flow LTE Turbo implementation trade-off
Same device and operating point. Better area and clock do not erase the 684-cycle latency cost.

Verification had to survive a deliberately bad design

The evidence chain moves through independent algorithm checks, RTL conformance, post-route implementation and board playback. The conformance record contains 153 complete runs across nine operating points. Cross-validation covers nine frame cases and 16,136 information bits. Board playback covers five legal block sizes from K=40 to K=6144, 8,883 input words and 8,848 decisions.

LTE Turbo decoder verification evidence chain
Each layer targets a different class of error. A green board record is paired with a negative control.

The current board record ran at 299.917 MHz for 3,003,668,343 cycles with zero mismatches. Then one RTL condition was deliberately mutated. That build produced 56,770,955 mismatches and failed. The control is important because a checker that never sees a known-bad design has not yet shown that it can reject one.

Current LTE Turbo decoder board record
Selected fields from the current machine-readable board evidence and its mutation control.

What is ready, and what still belongs in an evaluation

This case shows that AI-assisted RTL work can now produce a repeatable candidate-and-evidence package quickly enough to support a pipeline of communication IP, rather than a one-off demonstration. The package has a parameter boundary, a routed operating point, an explicit comparison, model and RTL records, and a board result with a refutation control.

It does not turn one device and one workload into a universal product guarantee. Customer delivery still needs the target device, interfaces, throughput, latency, channel model, licensing form and acceptance workload to be fixed. Those details determine which point should be generated, optimized and signed off.

The implementation figures and evaluation boundary are collected on the LTE Turbo Decoder product page.

Have a channel-decoding workload to turn into hardware?

Define the device, interfaces and acceptance workload. We will return a scoped candidate and the evidence needed to evaluate it.

Discuss an evaluation