GAUNTLET the release gate for AI inference hello@gauntletlab.com

The model you ship is not the model you validated.

Your model gets quantized and compiled before it reaches the device. Gauntlet measures exactly what that changed — and whether the compression or the silicon is to blame.

See a full report
INT8 · Ethos-U55-128 · Vela 5.2.0 · 500 eval inputs
FAILED exit code 1
1 fail · 3 not measured
placement.fallback_ops≤ 0 · 1FAIL
memory.flash_mb≤ 0.50 · 0.0146PASS
accuracy.answer_flip_rate≤ 0.020 · —N/M
latency.p99_ms≤ 15.0 · —N/M
Operators
9
Off the NPU
1

L2_NORMALIZATION fell to the CPU, and the build was blocked before a board existed. Three limits read N/M — not measured — because target execution wasn't available, and the report says which three and why rather than passing them.

How it works

Three columns. Every vendor tool has only one.

We run the same evaluation set three ways and compare them pairwise.

Measured drift · YOLOv8n INT8 · lower SNR means more damage
quantization · 26.2 dB silicon · 31.7 dB GOLDEN SCHEME TARGET
Golden

Your original FP32 model, run deterministically. Ground truth.

Scheme

The quantized model on reference kernels — the compression mathematics alone, no vendor code.

Target

The compiled artifact on the vendor toolchain, running on your board — in your enclosure, at your ambient. The only place thermal and bandwidth are real.

The loss was the compression, not the chip. Most of the distance from the original model is spent before the silicon is even involved — and that attribution is the one thing a chip vendor can never credibly tell you about its own silicon.
And the catch nobody surfaces

An operation left the accelerator and nothing told you.

Compilers silently reassign operations they can't place. The model still returns correct numbers, so every accuracy test passes — while the op runs on a CPU core and quietly eats the latency budget you sized the chip for. We read it straight out of the compiler's own output.

9 operators · 8 accelerated · 1 fell to the CPU. Vela's own words: Unsupported opType — an L2_NORMALIZATION, on Ethos-U55. The build was blocked with exit code 1 before a board existed. Read the full report

Get a qualification

Your model never leaves your network.

We ship you a pinned container. It runs on your bench, against your eval set, on your board — and the only thing that comes back is the report. No upload, no data room, no security review. We're working with a small number of teams shipping neural networks onto embedded silicon.

See a full report

Request a qualification

Tell us what you are shipping and onto what. We reply within two working days, and nothing you send here leaves our inbox.

Prefer email? hello@gauntletlab.com