Overview
The Sophgo BM1688 is the high-compute tier among Chinese edge TPUs. It is rated at 32 TOPS INT8 and up to 64 TOPS when quantised to INT4, all inside an 18W envelope that passive cooling can handle.
It contrasts with the Rockchip RK3588 in a straightforward way: the RK3588 wins on breadth and multimedia, while the BM1688 concentrates on high-compute and high-concurrency inference.
Key Features
- Compute tier: 32 TOPS at INT8 and up to 64 TOPS at INT4, positioning it against the NVIDIA Jetson Orin NX class.
- On-device large models: Native support for 7B-class models such as Llama3 and Qwen2 at the edge, with efficient INT4 quantisation. Measured throughput reaches 60 to 65 tokens per second on 7B models, close to a cloud experience, with multimodal image and text input supported.
- Video concurrency: Designed for multi-stream video analysis, supporting 32 parallel video streams, with image classification reaching 1,350 FPS on ResNet18.
- Power and thermals: Full-load power stays under 18W, allowing fanless passive cooling, and the industrial wide-temperature version covers -20°C to +60°C for outdoor and factory-floor deployment.
- Expansion and deployment: PCIe expansion with multi-card scaling up to four cards for a combined 128 TOPS. The toolchain is BMNNSDK2 with mainstream framework support and Docker-based deployment.
Use Cases
- On-device large models where local chat and multimodal inference need a usable 7B class model
- Heavy video analytics with 16 or more concurrent streams on a single unit
- Industrial and outdoor sites needing wide temperature, fanless operation and multi-card scaling
- Private inference deployments where data cannot leave the internal network
Pros
- A strong compute density at 32 TOPS INT8 within 18W
- Runs 7B class models on device with efficient INT4 quantisation
- Multi-card scaling up to 128 TOPS gives a clear growth path
- Industrial temperature range and fanless design suit field deployment
Pricing
Sold as chips and complete units. In 2026 BM1688 boxes sit roughly in the CNY 3,000 to 5,000 range, a flagship position suited to budget-conscious projects that still need real compute and value a domestic supply chain.
Summary
The BM1688 fits scenarios where enough compute, a low price and domestic sourcing all have to hold at once. 60 tokens per second on a 7B model at the edge is already usable, and multi-card scaling leaves room to grow. One thing to keep in mind is its positioning: the CPU is eight Cortex-A53 cores and the design is built around dedicated acceleration, so general-purpose compute is not the focus. If a project leans heavily on multimedia encoding and decoding, it is worth shortlisting the Rockchip RK3588 alongside it; the BM1688 suits deployments that prize compute density per board and multi-card expansion.
Version History
- BM1688 in production (2026): 32 TOPS INT8 and 64 TOPS INT4, supporting 7B-class on-device inference and up to four cards at 128 TOPS