BEVFormer-Tiny on Chimera: Bird's-Eye-View Perception in One GPNPU Kernel

BEVFormer-Tiny runs its neural-network stages in one Chimera GPNPU kernel. Six camera inputs, 16.1 million cycles per frame in the four-core instruction-set simulator. Here’s how ChiPy brings the network together.

AI models evolve faster than silicon, introducing operations that an accelerator may never have been designed to run. When that happens, peak TOPS tell you little: the deployment depends on whether the hardware can execute the whole network. BEVFormer brings that challenge into focus with attention that learns where to sample features, combining dense computation with data-dependent reads and interpolation.

The BEVFormer-Tiny demo in Chimera SDK 26.09 compiles the backbone, encoder, decoder, and prediction heads, including attention, into one GPNPU kernel and runs it across four GPNPU cores.

Chimera’s programmable GPNPU lets us implement sampling in C++ alongside the network’s dense computation, keeping attention on the same processor without a CPU or DSP handoff.

On a four-core GPNPU in the instruction-set simulator, the optimized build that will ship in SDK 26.10 processes one six-camera frame in 16.1M cycles: 9.5 ms at 1.7 GHz.

BEVFormer-Tiny INT8 detections from the four-core GPNPU instruction-set simulator, shown in red across six camera views and in BEV, with score ≥ 0.3.

Why Sampling Changes the Deployment

BEVFormer builds a bird's-eye-view feature map from six surround cameras, combining current images with information carried forward from earlier frames. Its spatial cross-attention samples image features around reference points projected into the cameras. Temporal self-attention samples current and historical BEV features after ego-motion alignment. A Deformable DETR-style decoder produces the features used to predict 3D boxes and velocity.

We used the Tiny configuration: ResNet-50, a 50×50 BEV grid, three encoder layers, one image-feature level, and 900 object queries.

Deformable attention predicts offsets and weights for each query. Those offsets become runtime read addresses; fractional positions require bilinear interpolation. The implementation has to combine dense projections with data-dependent gathers, coordinate arithmetic, boundary handling, and weighted sums.

Some geometry can be prepared ahead of time. For a fixed camera rig, image shape, and BEV grid, projection masks can be precomputed. The learned offsets still change with the input, so the sampling itself remains data-dependent. The official encoder implementation shows that distinction.

An export format doesn't settle how those operations execute. ONNX has GridSample, but a backend still needs an implementation for the exported or fused attention workload. In this demo, the backbone comes through ONNX and the encoder, decoder, and prediction heads use the Chimera Compute Library (CCL). ChiPy joins them into the compiled network.

How We Composed the Network with ChiPy

ChiPy is the Chimera SDK’s Python interface for composing neural-network modules and compiling them for the GPNPU. The implementation uses it to connect the ONNX backbone with encoder, decoder, and prediction modules defined in Python. Their forward methods make the operation boundaries explicit, while CCL supplies the GPNPU implementations. This lets us keep the network's structure visible in Python and compile the assembled chain into one kernel.

That is useful when attention would otherwise become a complicated web of operations in an exported graph. Developers can test sub-chains, swap in reference implementations, and isolate an issue at a module boundary before compiling for the GPNPU.

Calibration follows the same chain. A CPU forward pass records activation ranges at calibration points, which the workflow uses to determine fixed-point fractional-bit settings. The assembled chain also runs on the CPU as a reference for checking GPNPU execution.

For this demo, we reimplemented the model's forward logic and extracted the checkpoint weights. ChiPy gave us a defined workflow to assemble the modules, check their outputs, calibrate, and compile the network for the GPNPU. The SDK notebook walks through it.

16.1M Cycles per Frame on Four Cores

On the four-core ISS, one frame through the neural-network kernel takes 16.1M cycles, or 9.5 ms at 1.7 GHz. The demo compares its detection outputs with the ChiPy CPU reference on the same test scene.

GPNPU ISS detections (red) overlaid on the ChiPy CPU reference (blue), using the same test scene and score ≥ 0.3. This compares execution of the same pipeline; it is not a full-dataset accuracy evaluation.

Metric Chimera GPNPU
Model BEVFormer-Tiny
Input tensor 6 × 3 × 480 × 800
BEV grid 50×50
Precision INT8 quantized model
Configuration Four QC-U cores; 16 MACs/PE; 8 MB OCM; 1.7 GHz; 128 GB/s external bandwidth
Execution Instruction-set simulator (ISS)
Kernel cycles per frame 16,131,580
Simulated kernel latency at 1.7 GHz 9.5 ms
Postprocessing Host

Measurement note: cycle counts are full-frame ISS wall-clock for the neural-network kernel, from start and end markers on all four cores, using the notebook's hardware settings.

The SDK 26.09 demo notebook ships the kernels as first released. Since then, we have spent time optimizing them; the 16.1M-cycle figure comes from those optimized kernels, which will ship in SDK 26.10 and produce the same outputs.

The one-kernel result covers the backbone, encoder, decoder, and prediction heads. Sigmoid, box denormalization, top-k, and thresholding run on the host. This establishes compilation and simulated execution of the network, with a test-scene comparison against the reference pipeline. Full nuScenes accuracy, power, and camera-to-box timing remain to be established.

The Next Model Changes the Workload

BEVFormer and BEVFusion appeared in 2022; StreamPETR and SparseBEV followed in 2023. Sampling, temporal state, and sensor fusion keep changing while an automotive SoC takes years to develop.

The research keeps changing the workload. Programmable kernels give existing silicon a path to support new operations.

BEVFormer-Tiny is one more network running on Chimera. When the next architecture brings a new operation, we can support it the same way: write its kernel, compose it with the rest of the network in ChiPy, and validate it against the CPU reference, all on the GPNPU already designed into the SoC.

Run the BEVFormer-Tiny demo to inspect the implementation and reproduce the ISS workflow. Contact Quadric to discuss your next model's requirements.


Explore Quadric IP:


×
Semiconductor IP