At 19:30 UTC on October 6, an openTPU thread had 144 Hacker News points and 143 comments, nearly one reply for every point. The argument started with an unusually tangible artifact: an Apache-2.0 accelerator stack that its author says runs ten modern models on a Kintex-7 FPGA. It also carries a harder claim in its first sentence, that the accelerator was "developed by AI." The live discussion is a community-interest signal, not proof of either claim. The repository is where the useful evidence sits.
For developers, the interesting part is the boundary between those claims. openTPU exposes SystemVerilog, an instruction set, a bit-exact simulator, a small kernel language, a compiler, host tools and a browser profiler in one repository. Its AI agents can propose and implement changes, but they do not decide what counts as correct. A fixed evaluation pipeline does that job. That distinction makes the project more instructive than the familiar headline about AI building the hardware that runs AI.
A full stack on one older FPGA
The openTPU repository targets an Inspur YPCB-00338 PCIe card built around Xilinx's Kintex-7 xc7k480t and 4 GiB of DDR3 across two channels. The design uses a four-column systolic matrix unit, a vector unit for FP32 operations and explicit DMA instructions. There is no cache or hidden scheduler. Programs say where data moves, which lets the profiler account for cycles without guessing what a hardware scheduler did.
That simple structure reaches from Python to wires. A kernel decorated with @ol.jit becomes instructions made of eight 32-bit words. The Python ISA simulator and the SystemVerilog RTL are supposed to produce the same bits, and the same program can then run as a Vivado bitstream on the card. The project's otpu-lens tool records runs from the simulator, RTL or physical board and displays unit activity, memory waits and instruction timing in a browser.
The author's published measurements are specific enough to inspect. With 4-bit weights and an int8 language-model head, LFM2.5-230M decoded at 82.1 tokens per second wall-clock, Qwen3-0.6B at 30.7, and Qwen3.5-0.8B at 23.3. Larger dense models were slower: Phi-4-mini reached 6.55 tokens per second and Qwen3.5-4B reached 5.87. These are project results from one board and test method, not comparisons against current GPUs or CPUs.
The measurement notes also explain what each number includes. Decode tests generate 64 greedy tokens after a 512-token prompt. "Device" time counts accelerator cycles, while "wall" time adds the host. Every listed configuration is reported to match the simulator token for token. That scope matters because accelerator benchmarks can look better when host work, model loading or differing output choices disappear from the clock. openTPU at least labels the boundary.
Memory bandwidth sets the pace
The accelerator's ceiling is visible in its own table. Dense-model decode consumes between 82 and 94 percent of the board's stated 17.1 GB/s DDR3 peak. A newer build pushed prefill speed up by 1.3 to 2 times, depending on the model, while decode stayed within 2.3 percent of the previous image. The repository attributes that gap to memory: adding matrix throughput helps prompt processing, but token-by-token generation is already waiting on weights from DRAM.
Quantization buys speed by reducing those reads. openTPU's 4-bit format uses FP4 values with two levels of block scaling, for 4.25 bits per weight, while keeping the output head in int8. The project reports roughly one-third fewer bytes per token and decode gains of 40 percent for Qwen3.5 and 45 percent for Qwen3 and LFM2. It also reports a perplexity cost in its model notes. That last detail prevents a bandwidth win from masquerading as a free quality win.
Models larger than the card's memory take a different path. The board keeps a set of experts in DRAM and streams missing mixture-of-experts weights from host storage. In the author's October 1 run, LFM2.5-8B-A1B decoded at 10.6 tokens per second with a 98.5 percent expert-slot hit rate. Qwen3.5-35B-A3B managed 3.95 tokens per second, with a 62 percent hit rate and 153 MB transferred per token over PCIe. Both results are reported as bit-for-bit matches with the simulator.
None of this turns the Kintex-7 card into a general GPU replacement. The repository does not publish a controlled tokens-per-watt or price-performance comparison against modern alternatives. Even its board power figures are Vivado estimates marked low confidence because they lack switching activity. The narrower achievement is still real: several model families pass through one inspectable compiler, simulator and FPGA implementation.
What the agents are allowed to change
The phrase "developed by AI" becomes less mysterious inside the repository's tournament runtime. It defines three agent roles: a hypothesis writer, an implementer and a scribe. Claude Opus 5.5 is the default for those roles in the checked-in configuration, while Codex is also supported. The agents receive a component description, current synthesis metrics, a critical path report, recent outcomes and a strict list of files they may change.
Their authority is deliberately small. For a matrix-unit round, for example, the hypothesis prompt asks for one change that improves area or clock frequency without changing the computed result. The implementation agent works in an isolated Git worktree. Network tools are disabled for the Claude path, shell commands are limited, and a path check rejects edits outside the component's allowed files. The model can edit a candidate design and tests. It cannot rewrite the grader that decides whether the design survives.
The checked-in JSONL logs provide more than a polished success story. They record prompts, model settings, token use, cost, diffs, synthesis metrics and outcomes. One accepted matrix-unit round registered a scale-factor selection a stage earlier; the log reports a 3.8 percent estimated clock gain, from 122 to 127 MHz, with no area increase. Another cut dead lanes that the synthesizer could not prove dead, reducing the reported area-equivalent metric by 2.9 percent. These are narrow engineering changes attached to patches and gates, which makes them reviewable.
The predecessor project, auto-arch-tournament, makes the selection pressure even clearer. Its published run tested 73 hypotheses and accepted ten. Sixty-three failed its verifier. Each candidate passed through Verilator lint, synthesis, CoreMark, co-simulation against a Python instruction-set simulator, formal RISC-V checks and three FPGA place-and-route seeds. The agent supplied possible edits. Fixed constraints defined correctness, protected the evaluator and chose the fitness function.
That is the project's transferable idea. Agent-written code gets more credible when the surrounding system can cheaply reject plausible mistakes. Hardware makes the lesson stark because a syntactically valid SystemVerilog patch can still alter rounding, break a handshake, miss timing or consume too many lookup tables. openTPU asks the model for a hypothesis, then demands bit-exact behavior and measured improvement before accepting the patch.
Reproduction is possible, with limits
You can inspect much of the stack without the target card. The repository says the Python path runs on a laptop, and its quick start installs the package, PyTorch, Transformers and pytest before invoking the ISA simulator:
pip install -e .
pip install pytest torch transformers
python3 -m pytest -q
hf download LiquidAI/LFM2.5-230M --local-dir models/LFM2.5-230M
otpu-chat --model lfm2 --backend isa
Physical reproduction asks much more. The board guide calls for the specific Kintex-7 card, Vivado 2026.1 and a license that covers the xc7k480t. The free edition does not. AMD's 30-day evaluation license is an option. A normal bitstream build takes an estimated 1.5 to 3 hours, and the guide documents JTAG programming, PCIe setup, timing reports and memory calibration. That is enough detail to try, though the project page does not link to an independent board reproduction.
The next evidence should come from outside the author's machine. A useful reproduction would confirm token equality and wall-clock rates on another card, publish measured power, and compare the same quantized weights with a CPU or GPU baseline. The agent story needs one more accounting layer too: which subsystems began as human designs, which patches came from tournament agents, and where maintainers revised or rejected model output. Until then, openTPU is best read as an auditable accelerator and a disciplined experiment in AI-assisted hardware work. Its verifier, more than its opening slogan, is the part worth copying.