OpenTPU – An open-source AI accelerator, developed by AI
Article URL: https://github.com/FeSens/openTPU Comments URL: https://news.ycombinator.com/item?id=49980715 Points: 210 # Comments: 275
openTPU is an open-source AI accelerator developed by AI. It emerged from the auto-arch-tournament, questioning how far AI agents can go in hardware design and if they can construct the chip that runs their own inference. The accelerator's entire codebase is housed in a single monorepo, encompassing hardware design in SystemVerilog, an instruction set, a bit-exact simulator, a kernel language and its compiler, along with host software that operates a PCIe card.
Designed to run ten modern models with genuine weights on an Inspur YPCB-00338 card (Xilinx Kintex-7 xc7k480t, two DDR3 channels), the hardware delivers consistent token results with the simulator. Measured performance reveals that the initial three models showcased improved decoding times on September 29, 2026, with Build B executing LFM2-2.6B, SmolLM3-3B, and Phi-4-mini models approximately 8-9% faster than previous versions while maintaining 91-94% efficiency of DRAM peak.
Notably, the Qwen3.5-4B model exhibits an int8 image exceeding 4 GiB. The card systematically routes each token and computes every expert, storing them in per-layer slots within its DRAM. To optimize efficiency, the card streamlines experts from host storage when necessary, maintaining a robust data flow. The accelerator's design prioritizes simplicity, utilizing a sequencer to issue one instruction per cycle to various units.
These units include DMA for data movement, a matrix unit for int8 weight multiplication, a vector unit for FP32 mathematical operations, and a quantizer to revert results back to int8. The ISA, simulator, compiler, and profiler all reside within one repository, ensuring transparency and accessibility. The card's JTAG interface facilitates loading the bitstream, and users can initiate operations with sudo otpu-setup and otpu-chat --backend board.
Contributions, including issues and pull requests, are welcomed, although familiarity with Python and Verilator is required. Changes to the instruction set, simulator, or RTL must maintain passing criteria using python3 -m pytest -q. Performance claims must be substantiated with detailed measurement techniques.
Written by urgent.news from Hacker News Best's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.