INT8 Quantization & Pruning
Calibrated post-training quantization and structured pruning shrink models up to 4× with minimal accuracy loss.
Qovox compiles your trained ONNX or PyTorch model into lean, zero-dependency C++ — built for STM32, ESP32, ARM Cortex-M, RISC-V and automotive ECUs.
3 compiles a month. No card required.
# resnet8.onnx · graph summary input : float32[1,3,32,32] Conv → BatchNorm → Relu Conv → BatchNorm → Relu Add (residual) GlobalAveragePool Gemm (10 classes) output : float32[1,10] params: 78K · fp32 · 312 KB
import torch model = ResNet8().eval() x = torch.randn(1, 3, 32, 32) torch.onnx.export( model, x, "resnet8.onnx", opset_version=17) # then: # qovox compile resnet8.onnx \ # --target cortex-m4 --int8
// generated by Qovox · no libc++, no heap #pragma once #include <stdint.h> namespace qovox { constexpr size_t ARENA = 9216; void infer(const int8_t* in, int8_t* out); } // static arena · fused conv+bn+relu
Runs on your chipset · Reads your framework
Every byte of RAM and flash counts on a microcontroller. Qovox compiles it away.
Calibrated post-training quantization and structured pruning shrink models up to 4× with minimal accuracy loss.
Plain C++ header and weights. No runtime, no heap, no vendor lock-in.
Deterministic, static-memory output suited to ECU and CAN-bus pipelines.
Fused layers and SIMD-friendly loops keep small vision and signal models in real-time budgets.
Same model, same board, same accuracy.
Illustrative sample (ResNet-8, Cortex-M4). Shorter bars are better.
Size and latency before and after Qovox INT8 compilation. Filter by task or search by name.
| Model | Category | FP32 → INT8 size | Compression | FP32 → Qovox WASM latency | Speedup | Actions |
|---|
Illustrative sample figures, not measured results. Latency is single-inference time in the browser (WASM).
Three steps from a trained model to code you can build.
Drop in an ONNX, PyTorch, TensorFlow or GGUF model file.
Qovox fuses layers, quantizes weights and emits C++.
Download the source and build it anywhere with CMake.
Pick a sample model. ONNX models run locally with ONNX Runtime Web (nothing is uploaded); other formats show a simulated Qovox compiler pipeline.
The model loads when this section comes into view.
Ctrl/⌘ + Enter to run.
No image selected? A random sample input is used.
Select a model and press “Run Model”.