Repository navigation
Clarify zero copy semantics #56
Description
Activity
I just want to add that prost supports bytes which is meant to provide zero-copy decoding for
bytefields in protobuf messages. I think this repo should try benchmarking with that too for a more fair benchmark.For example, running the prost benchmarks locally with `bytes` generation enabled
➜ buffa git:(main) ✗ cd benchmarks/prost && cargo bench Compiling bench-prost v0.1.0 (/home/dezyh/code/oss/buffa/benchmarks/prost) Finished `bench` profile [optimized] target(s) in 2.75s Running benches/protobuf.rs (target/release/deps/protobuf-d4f9fd6e9a32ec34) Gnuplot not found, using plotters backend prost/api_response/decode time: [491.21 ns 494.92 ns 500.07 ns] thrpt: [12.003 GiB/s 12.128 GiB/s 12.220 GiB/s] change: time: [-93.983% -93.962% -93.935%] (p = 0.00 < 0.05) thrpt: [+1548.8% +1556.1% +1562.1%] Performance has improved. Found 4 outliers among 100 measurements (4.00%) 1 (1.00%) low severe 3 (3.00%) high severe prost/api_response/merge time: [178.85 ns 179.17 ns 179.54 ns] thrpt: [33.431 GiB/s 33.502 GiB/s 33.561 GiB/s] change: time: [-97.045% -97.034% -97.023%] (p = 0.00 < 0.05) thrpt: [+3259.3% +3271.8% +3284.5%] Performance has improved. Found 4 outliers among 100 measurements (4.00%) 2 (2.00%) high mild 2 (2.00%) high severe prost/api_response/encode time: [450.67 ns 453.19 ns 455.99 ns] thrpt: [13.163 GiB/s 13.245 GiB/s 13.319 GiB/s] change: time: [-84.151% -84.072% -83.990%] (p = 0.00 < 0.05) thrpt: [+524.61% +527.82% +530.94%] Performance has improved. Found 9 outliers among 100 measurements (9.00%) 4 (4.00%) low mild 4 (4.00%) high mild 1 (1.00%) high severe prost/api_response/encoded_len time: [125.05 ns 127.13 ns 129.01 ns] thrpt: [46.525 GiB/s 47.215 GiB/s 47.998 GiB/s] change: time: [-50.197% -49.012% -47.903%] (p = 0.00 < 0.05) thrpt: [+91.950% +96.125% +100.79%] Performance has improved. Found 3 outliers among 100 measurements (3.00%) 3 (3.00%) high mild prost/api_response/json_encode time: [16.983 µs 17.030 µs 17.084 µs] thrpt: [619.70 MiB/s 621.65 MiB/s 623.37 MiB/s] change: time: [+12.063% +12.457% +12.828%] (p = 0.00 < 0.05) thrpt: [-11.369% -11.077% -10.765%] Performance has regressed. Found 9 outliers among 100 measurements (9.00%) 8 (8.00%) high mild 1 (1.00%) high severe prost/api_response/json_decode time: [43.373 µs 43.540 µs 43.773 µs] thrpt: [241.85 MiB/s 243.15 MiB/s 244.09 MiB/s] change: time: [+17.023% +18.008% +18.909%] (p = 0.00 < 0.05) thrpt: [-15.902% -15.260% -14.546%] Performance has regressed. Found 7 outliers among 100 measurements (7.00%) 4 (4.00%) high mild 3 (3.00%) high severe prost/log_record/decode time: [915.73 ns 935.83 ns 958.23 ns] thrpt: [30.964 GiB/s 31.706 GiB/s 32.401 GiB/s] change: time: [-97.799% -97.760% -97.719%] (p = 0.00 < 0.05) thrpt: [+4283.6% +4364.0% +4444.0%] Performance has improved. Found 27 outliers among 100 measurements (27.00%) 15 (15.00%) low severe 2 (2.00%) low mild 3 (3.00%) high mild 7 (7.00%) high severe prost/log_record/merge time: [351.96 ns 370.86 ns 392.27 ns] thrpt: [75.640 GiB/s 80.005 GiB/s 84.303 GiB/s] change: time: [-98.960% -98.933% -98.899%] (p = 0.00 < 0.05) thrpt: [+8979.5% +9271.5% +9514.4%] Performance has improved. Found 22 outliers among 100 measurements (22.00%) 1 (1.00%) low mild 4 (4.00%) high mild 17 (17.00%) high severe prost/log_record/encode time: [469.78 ns 481.27 ns 496.41 ns] thrpt: [59.771 GiB/s 61.652 GiB/s 63.159 GiB/s] change: time: [-93.923% -93.814% -93.685%] (p = 0.00 < 0.05) thrpt: [+1483.5% +1516.6% +1545.6%] Performance has improved. Found 1 outliers among 100 measurements (1.00%) 1 (1.00%) high mild prost/log_record/encoded_len time: [169.94 ns 170.55 ns 171.20 ns] thrpt: [173.31 GiB/s 173.97 GiB/s 174.60 GiB/s] change: time: [-83.292% -83.161% -83.070%] (p = 0.00 < 0.05) thrpt: [+490.68% +493.87% +498.52%] Performance has improved. Found 3 outliers among 100 measurements (3.00%) 2 (2.00%) low mild 1 (1.00%) high mild prost/log_record/json_encode time: [36.663 µs 37.048 µs 37.461 µs] thrpt: [996.24 MiB/s 1007.3 MiB/s 1017.9 MiB/s] change: time: [+2.2195% +3.4337% +4.7411%] (p = 0.00 < 0.05) thrpt: [-4.5265% -3.3197% -2.1713%] Performance has regressed. Found 20 outliers among 100 measurements (20.00%) 8 (8.00%) high mild 12 (12.00%) high severe prost/log_record/json_decode time: [57.339 µs 57.955 µs 58.672 µs] thrpt: [636.08 MiB/s 643.95 MiB/s 650.87 MiB/s] change: time: [+2.2690% +3.1342% +4.0758%] (p = 0.00 < 0.05) thrpt: [-3.9162% -3.0390% -2.2187%] Performance has regressed. Found 14 outliers among 100 measurements (14.00%) 1 (1.00%) high mild 13 (13.00%) high severe prost/analytics_event/decode time: [546.48 ns 547.04 ns 547.65 ns] thrpt: [268.33 GiB/s 268.63 GiB/s 268.90 GiB/s] change: time: [-99.910% -99.909% -99.909%] (p = 0.00 < 0.05) thrpt: [+109438% +110187% +111325%] Performance has improved. Found 8 outliers among 100 measurements (8.00%) 1 (1.00%) low severe 2 (2.00%) low mild 4 (4.00%) high mild 1 (1.00%) high severe prost/analytics_event/merge time: [251.98 ns 252.76 ns 253.75 ns] thrpt: [579.12 GiB/s 581.40 GiB/s 583.19 GiB/s] change: time: [-99.956% -99.955% -99.955%] (p = 0.00 < 0.05) thrpt: [+220944% +223261% +225890%] Performance has improved. prost/analytics_event/encode time: [387.63 ns 388.01 ns 388.57 ns] thrpt: [378.19 GiB/s 378.73 GiB/s 379.10 GiB/s] change: time: [-99.889% -99.889% -99.887%] (p = 0.00 < 0.05) thrpt: [+88596% +89594% +90223%] Performance has improved. Found 2 outliers among 100 measurements (2.00%) 1 (1.00%) high mild 1 (1.00%) high severe prost/analytics_event/encoded_len time: [95.358 ns 95.426 ns 95.494 ns] thrpt: [1538.8 GiB/s 1539.9 GiB/s 1541.0 GiB/s] change: time: [-99.795% -99.794% -99.792%] (p = 0.00 < 0.05) thrpt: [+48028% +48367% +48751%] Performance has improved. Found 6 outliers among 100 measurements (6.00%) 1 (1.00%) low severe 2 (2.00%) low mild 3 (3.00%) high mild prost/analytics_event/json_encode time: [373.44 µs 373.88 µs 374.33 µs] thrpt: [797.02 MiB/s 797.99 MiB/s 798.94 MiB/s] change: time: [-0.6800% -0.5036% -0.3330%] (p = 0.00 < 0.05) thrpt: [+0.3341% +0.5062% +0.6847%] Change within noise threshold. Found 1 outliers among 100 measurements (1.00%) 1 (1.00%) high severe Benchmarking prost/analytics_event/json_decode: Warming up for 3.0000 s Warning: Unable to complete 100 samples in 5.0s. You may wish to increase target time to 6.5s, enable flat sampling, or reduce sample count to 60. prost/analytics_event/json_decode time: [1.2494 ms 1.2500 ms 1.2506 ms] thrpt: [238.56 MiB/s 238.68 MiB/s 238.79 MiB/s] change: time: [-1.2606% -1.1312% -0.9973%] (p = 0.00 < 0.05) thrpt: [+1.0074% +1.1442% +1.2767%] Change within noise threshold. Found 6 outliers among 100 measurements (6.00%) 4 (4.00%) high mild 2 (2.00%) high severe prost/google_message1_proto3/decode time: [16.339 ns 16.344 ns 16.350 ns] thrpt: [12.988 GiB/s 12.992 GiB/s 12.996 GiB/s] change: time: [-92.280% -92.198% -92.130%] (p = 0.00 < 0.05) thrpt: [+1170.6% +1181.7% +1195.4%] Performance has improved. Found 17 outliers among 100 measurements (17.00%) 4 (4.00%) low severe 1 (1.00%) low mild 2 (2.00%) high mild 10 (10.00%) high severe prost/google_message1_proto3/merge time: [6.8130 ns 6.8173 ns 6.8219 ns] thrpt: [31.126 GiB/s 31.147 GiB/s 31.167 GiB/s] change: time: [-95.580% -95.571% -95.561%] (p = 0.00 < 0.05) thrpt: [+2152.7% +2158.0% +2162.7%] Performance has improved. Found 11 outliers among 100 measurements (11.00%) 10 (10.00%) high mild 1 (1.00%) high severe prost/google_message1_proto3/encode time: [21.118 ns 21.165 ns 21.214 ns] thrpt: [10.010 GiB/s 10.033 GiB/s 10.055 GiB/s] change: time: [-80.394% -80.310% -80.230%] (p = 0.00 < 0.05) thrpt: [+405.82% +407.87% +410.06%] Performance has improved. prost/google_message1_proto3/encoded_len time: [9.5583 ns 9.7381 ns 9.9516 ns] thrpt: [21.337 GiB/s 21.805 GiB/s 22.215 GiB/s] change: time: [-40.450% -39.920% -39.117%] (p = 0.00 < 0.05) thrpt: [+64.249% +66.444% +67.927%] Performance has improved. Found 13 outliers among 100 measurements (13.00%) 1 (1.00%) low severe 2 (2.00%) low mild 1 (1.00%) high mild 9 (9.00%) high severe prost/google_message1_proto3/json_encode time: [544.20 ns 559.85 ns 574.51 ns] thrpt: [683.91 MiB/s 701.82 MiB/s 722.01 MiB/s] change: time: [+3.5656% +6.4151% +9.4428%] (p = 0.00 < 0.05) thrpt: [-8.6281% -6.0284% -3.4428%] Performance has regressed. Found 1 outliers among 100 measurements (1.00%) 1 (1.00%) high severe prost/google_message1_proto3/json_decode time: [1.6379 µs 1.6476 µs 1.6602 µs] thrpt: [236.67 MiB/s 238.48 MiB/s 239.89 MiB/s] change: time: [+0.9807% +1.5064% +2.0611%] (p = 0.00 < 0.05) thrpt: [-2.0195% -1.4840% -0.9711%] Change within noise threshold. Found 8 outliers among 100 measurements (8.00%) 4 (4.00%) high mild 4 (4.00%) high severe@brancz "zero copy" in the README is shorthand for the payload bytes of length-delimited fields are not copied from the input buffer into owned Rust types during decode. Varints, tags, and the fixed-width scalars are read, not copied in the memcpy sense. If that distinction would be clearer in the README we can tighten the wording — the claim we're making is narrower than "nothing is read from the input," and we don't want it to read as if it were.
What a view does avoid
A view is a borrow over the input buffer. For every length-delimited field, a view stores a slice into the original bytes:
stringfield →&'a str(UTF-8 validated once, no allocation)bytesfield →&'a [u8]- Nested
message→SubMessageView<'a>(another borrow, no recursion into owned types) repeated T(packed or otherwise) →RepeatedView<'a, T>— an iterator-style wrapper over a slice of the input; elements are lazily materialized on.iter()map<K, V>→MapView<'a, K, V>— same idea; the map is walked, not built
The result is that
decode_viewon a typical message does zero heap allocations, makes one linear pass over the wire format to record per-field slice offsets, and defers all per-element decoding to the caller. TheMessageViewtrait is the contract.What a view cannot avoid
Protobuf's wire format is not a memory layout, so some re-encodings are inevitable:
- Varint + tag reads. Every field needs its tag and (for length-delimited fields) its length decoded to find the next field. That is a per-field constant cost independent of field size — typically 1 or 2 bytes of varint. buffa's view decoder does exactly this linear scan once, records a small per-field descriptor (wire offset + length for length-delimited types; or the decoded value for fixed-width scalars), and stops. It does not copy payload bytes.
- Scalars that are not length-delimited (
int32,fixed64,bool, etc.) are decoded eagerly during the scan, because there's nothing meaningful to borrow — they are already their target Rust type. Forfixed32/fixed64/sfixed*/float/double, this is a small-constant byte swap on big-endian hosts and a pointer-read on little-endian; there is no "zero copy" representation for these that would also be correct cross-platform. - UTF-8 validation for
stringfields. A view holds&'a str, not&'a [u8], so we check once that the bytes are well-formed UTF-8. That is required by the protobuf spec and by Rust'sstrinvariant — no encoding lets us skip it without giving up one of the two. - Unknown-field scanning. Views preserve unknown fields the same way owned messages do; the skipping pass reads tag+length for each unknown field. Again, linear in field count, not in payload size.
The scan itself is eager (not lazy), because the per-field offset table is what makes subsequent field access O(1) without rescanning. A fully lazy design would push that cost onto the first field access instead, which is a different trade-off; buffa does not take it.
Re: prost's
bytesfeature: Fair point @dezyh — we should benchmark prost with itsbytes::Bytesfeature enabled, both for completeness and because it's the closest prost analogue to what a view does forbytesfields. We've now done that in PR #61. The newprost (bytes)benchmark variant appliesprost-build::Config::bytes(["."])(substitutingbytes::Bytesfor everybytesfield) and decodes frombytes::Bytesinput, which is what actually exercises prost's zero-copycopy_to_bytesslicing path. The benchmark suite also now includes a deliberately bytes-heavyMediaFramemessage (primarybytes body+repeated bytes chunks+map<string, bytes> attachments) so the feature has something to work with.Raw throughput (MiB/s, higher is better; Intel Xeon Platinum 8488C,
task bench-crossin Docker;prost (bytes)deltas are relative toprost,buffa (view)deltas relative tobuffa owned):Message prost prost (bytes) buffa owned buffa (view) ApiResponse 756 676 (−11%) 862 1,475 (+71%) LogRecord 712 676 (−5%) 722 1,984 (+175%) AnalyticsEvent 254 194 (−24%) 199 320 (+61%) GoogleMessage1 956 931 (−3%) 1,014 1,341 (+32%) MediaFrame 9,648 23,516 (+144%) 16,816 73,004 (+334%) Observations:
-
prost (bytes)tracks default prost within noise on the larger messages and is ~10% slower onApiResponse. The slow-down onApiResponseis real: the proto message is small (~128 bytes/payload) and per-message fixed cost dominates, and enablingbytesresults in extra work for each message decode. Specifically, the atomicfetch_add/fetch_subaround eachpayload.clone()plus theBytesAdapter-generic codepath run even if there are nobytesfields. This adds up to nontrivial overhead without any zero-copy payoff to offset it. On larger single messages that overhead is a smaller fraction of the total decode cost (which explains the results forLogRecord/GoogleMessage1), and with deeply nested small messages likeAnalyticsEvent, the overheads makebytessignificantly worse (~24% slower here). -
On
MediaFrame,prost (bytes)is ~2.4× faster than default prost. This is exactly the payoff the feature advertises — the primary body, the chunks, and the attachments are all sliced from the input buffer rather than copied into freshVec<u8>s. This overcomes the overheads seen for the other message types that don't contain bytes fields. -
buffa views remain ~3× faster than
prost (bytes)onMediaFrame. To explain where the remaining gap comes from, we ranperf staton a native--measurement-time 6run of each variant onMediaFramedecode. Sample results:variant GiB/s cyc/msg ins/msg L1-miss/msg IPC prost 9.1 3,929 10,470 113 2.67 buffa owned 16.6 2,101 5,679 91 2.70 prost (bytes) 21.9 1,632 4,625 5 2.83 buffa (view) 69.8 508 1,675 4 3.30 prost (bytes)closed the cache-miss gap — L1-dcache misses drop 113 → 5, which is the allocator traffic from cloningbytespayloads disappearing. But it still runs 2.8× more instructions per message than views: ~3k instructions per message is spent onStringallocation forframe_idandcontent_type, UTF-8 validation of those strings, theHashMap<String, Bytes>theattachmentsmap still builds (with hashed + heap-copied keys), and theVec<Bytes>thechunksfield still collects. Views skip all of it: the two strings stay&'a str,attachmentsis aMapView<'a, &'a str, &'a [u8]>that walks the input lazily on iter, andchunksis aRepeatedView<'a, &'a [u8]>.
So
prost (bytes)and buffa views are solving different (overlapping) problems:.bytes(["."])is the correct prost feature for "don't copy bytes payloads", and it delivers cleanly on that. Views go further by not building the owning Rust collections for strings, maps, and repeateds either. When those fields aren't in your schema, the two should converge; when they are, the gap is what we see above.- added a commit that references this issue
on Apr 23, 2026 I think this could be even clearer. It’s starting to make sense though, the point is that the input buffer never gets copied (hence zero copy), but it’s not zero-heap-allocations since the offset mask still needs to be computed.
- locked and limited conversation to collaborators
on Apr 24, 2026
It’s unclear from the readme to what extend zero copy actually holds true. By my understanding, protobuf uses varints to encode any integers, which I would consider not to be zero-copy since we need to read up to 10bytes per integer in order to know offsets of other elements. It would be great if the readme explained that behavior and also when this resolution happens. Lazily when requesting to read certain parts? Eargerly (doesn’t seem zero copy either as the tree of messages would need to eagerly need to create a mask of offsets)?