Skip to content

openrknn: fix FP16 output copy — bit-exact vendor match - #93

Merged
widgetii merged 1 commit into
masterfrom
openrknn/fix-fp16-output-copy
Apr 13, 2026
Merged

widgetii merged 1 commit into
masterfrom
openrknn/fix-fp16-output-copy

Conversation

@widgetii

Copy link
Copy Markdown
Owner

Summary

One-line fix: non-4D output path copied n_elems bytes instead of ti->size bytes. For FP16 3D tensors, this halved the output data (786432 bytes instead of 1572864).

Result

SmolVLM l0_mlp: bit-exact with vendor rknnlite2

  • 786432/786432 FP16 values match (cosine = 1.000000)
  • After [128,768,8] → [1024,768] detiling, output equals vendor byte-for-byte

Test plan

  • CI: 9/9 pass
  • l0_mlp: 100% FP16 exact match with vendor

🤖 Generated with Claude Code

The non-4D output path copied n_elems bytes instead of ti->size bytes.
For FP16 3D tensors like SmolVLM [1,1024,768], n_elems=786432 but the
actual data size is 1,572,864 bytes (786432 × 2). This caused 50% of
the output to be zero.

Fix: use ti->size (byte count) instead of ti->n_elems (element count).

Result: SmolVLM l0_mlp output is now **bit-exact** with vendor rknnlite2
(786432/786432 FP16 values match, cosine=1.000000).

Also adds ORKNN_DUMP_ACT env var for activation BO post-run dump.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
@widgetii
widgetii merged commit 34212a2 into master Apr 13, 2026
7 checks passed
@widgetii
widgetii deleted the openrknn/fix-fp16-output-copy branch April 13, 2026 11:00
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant