Skip to content

Clarify zero copy semantics #56

Description

@brancz

It’s unclear from the readme to what extend zero copy actually holds true. By my understanding, protobuf uses varints to encode any integers, which I would consider not to be zero-copy since we need to read up to 10bytes per integer in order to know offsets of other elements. It would be great if the readme explained that behavior and also when this resolution happens. Lazily when requesting to read certain parts? Eargerly (doesn’t seem zero copy either as the tree of messages would need to eagerly need to create a mask of offsets)?

Activity

  1. dezyh commented on Apr 21, 2026

    @dezyh

    I just want to add that prost supports bytes which is meant to provide zero-copy decoding for byte fields in protobuf messages. I think this repo should try benchmarking with that too for a more fair benchmark.

    For example, running the prost benchmarks locally with `bytes` generation enabled
    ➜  buffa git:(main) ✗ cd benchmarks/prost && cargo bench
       Compiling bench-prost v0.1.0 (/home/dezyh/code/oss/buffa/benchmarks/prost)
        Finished `bench` profile [optimized] target(s) in 2.75s
         Running benches/protobuf.rs (target/release/deps/protobuf-d4f9fd6e9a32ec34)
    Gnuplot not found, using plotters backend
    prost/api_response/decode
                            time:   [491.21 ns 494.92 ns 500.07 ns]
                            thrpt:  [12.003 GiB/s 12.128 GiB/s 12.220 GiB/s]
                     change:
                            time:   [-93.983% -93.962% -93.935%] (p = 0.00 < 0.05)
                            thrpt:  [+1548.8% +1556.1% +1562.1%]
                            Performance has improved.
    Found 4 outliers among 100 measurements (4.00%)
      1 (1.00%) low severe
      3 (3.00%) high severe
    prost/api_response/merge
                            time:   [178.85 ns 179.17 ns 179.54 ns]
                            thrpt:  [33.431 GiB/s 33.502 GiB/s 33.561 GiB/s]
                     change:
                            time:   [-97.045% -97.034% -97.023%] (p = 0.00 < 0.05)
                            thrpt:  [+3259.3% +3271.8% +3284.5%]
                            Performance has improved.
    Found 4 outliers among 100 measurements (4.00%)
      2 (2.00%) high mild
      2 (2.00%) high severe
    prost/api_response/encode
                            time:   [450.67 ns 453.19 ns 455.99 ns]
                            thrpt:  [13.163 GiB/s 13.245 GiB/s 13.319 GiB/s]
                     change:
                            time:   [-84.151% -84.072% -83.990%] (p = 0.00 < 0.05)
                            thrpt:  [+524.61% +527.82% +530.94%]
                            Performance has improved.
    Found 9 outliers among 100 measurements (9.00%)
      4 (4.00%) low mild
      4 (4.00%) high mild
      1 (1.00%) high severe
    prost/api_response/encoded_len
                            time:   [125.05 ns 127.13 ns 129.01 ns]
                            thrpt:  [46.525 GiB/s 47.215 GiB/s 47.998 GiB/s]
                     change:
                            time:   [-50.197% -49.012% -47.903%] (p = 0.00 < 0.05)
                            thrpt:  [+91.950% +96.125% +100.79%]
                            Performance has improved.
    Found 3 outliers among 100 measurements (3.00%)
      3 (3.00%) high mild
    
    prost/api_response/json_encode
                            time:   [16.983 µs 17.030 µs 17.084 µs]
                            thrpt:  [619.70 MiB/s 621.65 MiB/s 623.37 MiB/s]
                     change:
                            time:   [+12.063% +12.457% +12.828%] (p = 0.00 < 0.05)
                            thrpt:  [-11.369% -11.077% -10.765%]
                            Performance has regressed.
    Found 9 outliers among 100 measurements (9.00%)
      8 (8.00%) high mild
      1 (1.00%) high severe
    prost/api_response/json_decode
                            time:   [43.373 µs 43.540 µs 43.773 µs]
                            thrpt:  [241.85 MiB/s 243.15 MiB/s 244.09 MiB/s]
                     change:
                            time:   [+17.023% +18.008% +18.909%] (p = 0.00 < 0.05)
                            thrpt:  [-15.902% -15.260% -14.546%]
                            Performance has regressed.
    Found 7 outliers among 100 measurements (7.00%)
      4 (4.00%) high mild
      3 (3.00%) high severe
    
    prost/log_record/decode time:   [915.73 ns 935.83 ns 958.23 ns]
                            thrpt:  [30.964 GiB/s 31.706 GiB/s 32.401 GiB/s]
                     change:
                            time:   [-97.799% -97.760% -97.719%] (p = 0.00 < 0.05)
                            thrpt:  [+4283.6% +4364.0% +4444.0%]
                            Performance has improved.
    Found 27 outliers among 100 measurements (27.00%)
      15 (15.00%) low severe
      2 (2.00%) low mild
      3 (3.00%) high mild
      7 (7.00%) high severe
    prost/log_record/merge  time:   [351.96 ns 370.86 ns 392.27 ns]
                            thrpt:  [75.640 GiB/s 80.005 GiB/s 84.303 GiB/s]
                     change:
                            time:   [-98.960% -98.933% -98.899%] (p = 0.00 < 0.05)
                            thrpt:  [+8979.5% +9271.5% +9514.4%]
                            Performance has improved.
    Found 22 outliers among 100 measurements (22.00%)
      1 (1.00%) low mild
      4 (4.00%) high mild
      17 (17.00%) high severe
    prost/log_record/encode time:   [469.78 ns 481.27 ns 496.41 ns]
                            thrpt:  [59.771 GiB/s 61.652 GiB/s 63.159 GiB/s]
                     change:
                            time:   [-93.923% -93.814% -93.685%] (p = 0.00 < 0.05)
                            thrpt:  [+1483.5% +1516.6% +1545.6%]
                            Performance has improved.
    Found 1 outliers among 100 measurements (1.00%)
      1 (1.00%) high mild
    prost/log_record/encoded_len
                            time:   [169.94 ns 170.55 ns 171.20 ns]
                            thrpt:  [173.31 GiB/s 173.97 GiB/s 174.60 GiB/s]
                     change:
                            time:   [-83.292% -83.161% -83.070%] (p = 0.00 < 0.05)
                            thrpt:  [+490.68% +493.87% +498.52%]
                            Performance has improved.
    Found 3 outliers among 100 measurements (3.00%)
      2 (2.00%) low mild
      1 (1.00%) high mild
    
    prost/log_record/json_encode
                            time:   [36.663 µs 37.048 µs 37.461 µs]
                            thrpt:  [996.24 MiB/s 1007.3 MiB/s 1017.9 MiB/s]
                     change:
                            time:   [+2.2195% +3.4337% +4.7411%] (p = 0.00 < 0.05)
                            thrpt:  [-4.5265% -3.3197% -2.1713%]
                            Performance has regressed.
    Found 20 outliers among 100 measurements (20.00%)
      8 (8.00%) high mild
      12 (12.00%) high severe
    prost/log_record/json_decode
                            time:   [57.339 µs 57.955 µs 58.672 µs]
                            thrpt:  [636.08 MiB/s 643.95 MiB/s 650.87 MiB/s]
                     change:
                            time:   [+2.2690% +3.1342% +4.0758%] (p = 0.00 < 0.05)
                            thrpt:  [-3.9162% -3.0390% -2.2187%]
                            Performance has regressed.
    Found 14 outliers among 100 measurements (14.00%)
      1 (1.00%) high mild
      13 (13.00%) high severe
    
    prost/analytics_event/decode
                            time:   [546.48 ns 547.04 ns 547.65 ns]
                            thrpt:  [268.33 GiB/s 268.63 GiB/s 268.90 GiB/s]
                     change:
                            time:   [-99.910% -99.909% -99.909%] (p = 0.00 < 0.05)
                            thrpt:  [+109438% +110187% +111325%]
                            Performance has improved.
    Found 8 outliers among 100 measurements (8.00%)
      1 (1.00%) low severe
      2 (2.00%) low mild
      4 (4.00%) high mild
      1 (1.00%) high severe
    prost/analytics_event/merge
                            time:   [251.98 ns 252.76 ns 253.75 ns]
                            thrpt:  [579.12 GiB/s 581.40 GiB/s 583.19 GiB/s]
                     change:
                            time:   [-99.956% -99.955% -99.955%] (p = 0.00 < 0.05)
                            thrpt:  [+220944% +223261% +225890%]
                            Performance has improved.
    prost/analytics_event/encode
                            time:   [387.63 ns 388.01 ns 388.57 ns]
                            thrpt:  [378.19 GiB/s 378.73 GiB/s 379.10 GiB/s]
                     change:
                            time:   [-99.889% -99.889% -99.887%] (p = 0.00 < 0.05)
                            thrpt:  [+88596% +89594% +90223%]
                            Performance has improved.
    Found 2 outliers among 100 measurements (2.00%)
      1 (1.00%) high mild
      1 (1.00%) high severe
    prost/analytics_event/encoded_len
                            time:   [95.358 ns 95.426 ns 95.494 ns]
                            thrpt:  [1538.8 GiB/s 1539.9 GiB/s 1541.0 GiB/s]
                     change:
                            time:   [-99.795% -99.794% -99.792%] (p = 0.00 < 0.05)
                            thrpt:  [+48028% +48367% +48751%]
                            Performance has improved.
    Found 6 outliers among 100 measurements (6.00%)
      1 (1.00%) low severe
      2 (2.00%) low mild
      3 (3.00%) high mild
    
    prost/analytics_event/json_encode
                            time:   [373.44 µs 373.88 µs 374.33 µs]
                            thrpt:  [797.02 MiB/s 797.99 MiB/s 798.94 MiB/s]
                     change:
                            time:   [-0.6800% -0.5036% -0.3330%] (p = 0.00 < 0.05)
                            thrpt:  [+0.3341% +0.5062% +0.6847%]
                            Change within noise threshold.
    Found 1 outliers among 100 measurements (1.00%)
      1 (1.00%) high severe
    Benchmarking prost/analytics_event/json_decode: Warming up for 3.0000 s
    Warning: Unable to complete 100 samples in 5.0s. You may wish to increase target time to 6.5s, enable flat sampling, or reduce sample count to 60.
    prost/analytics_event/json_decode
                            time:   [1.2494 ms 1.2500 ms 1.2506 ms]
                            thrpt:  [238.56 MiB/s 238.68 MiB/s 238.79 MiB/s]
                     change:
                            time:   [-1.2606% -1.1312% -0.9973%] (p = 0.00 < 0.05)
                            thrpt:  [+1.0074% +1.1442% +1.2767%]
                            Change within noise threshold.
    Found 6 outliers among 100 measurements (6.00%)
      4 (4.00%) high mild
      2 (2.00%) high severe
    
    prost/google_message1_proto3/decode
                            time:   [16.339 ns 16.344 ns 16.350 ns]
                            thrpt:  [12.988 GiB/s 12.992 GiB/s 12.996 GiB/s]
                     change:
                            time:   [-92.280% -92.198% -92.130%] (p = 0.00 < 0.05)
                            thrpt:  [+1170.6% +1181.7% +1195.4%]
                            Performance has improved.
    Found 17 outliers among 100 measurements (17.00%)
      4 (4.00%) low severe
      1 (1.00%) low mild
      2 (2.00%) high mild
      10 (10.00%) high severe
    prost/google_message1_proto3/merge
                            time:   [6.8130 ns 6.8173 ns 6.8219 ns]
                            thrpt:  [31.126 GiB/s 31.147 GiB/s 31.167 GiB/s]
                     change:
                            time:   [-95.580% -95.571% -95.561%] (p = 0.00 < 0.05)
                            thrpt:  [+2152.7% +2158.0% +2162.7%]
                            Performance has improved.
    Found 11 outliers among 100 measurements (11.00%)
      10 (10.00%) high mild
      1 (1.00%) high severe
    prost/google_message1_proto3/encode
                            time:   [21.118 ns 21.165 ns 21.214 ns]
                            thrpt:  [10.010 GiB/s 10.033 GiB/s 10.055 GiB/s]
                     change:
                            time:   [-80.394% -80.310% -80.230%] (p = 0.00 < 0.05)
                            thrpt:  [+405.82% +407.87% +410.06%]
                            Performance has improved.
    prost/google_message1_proto3/encoded_len
                            time:   [9.5583 ns 9.7381 ns 9.9516 ns]
                            thrpt:  [21.337 GiB/s 21.805 GiB/s 22.215 GiB/s]
                     change:
                            time:   [-40.450% -39.920% -39.117%] (p = 0.00 < 0.05)
                            thrpt:  [+64.249% +66.444% +67.927%]
                            Performance has improved.
    Found 13 outliers among 100 measurements (13.00%)
      1 (1.00%) low severe
      2 (2.00%) low mild
      1 (1.00%) high mild
      9 (9.00%) high severe
    
    prost/google_message1_proto3/json_encode
                            time:   [544.20 ns 559.85 ns 574.51 ns]
                            thrpt:  [683.91 MiB/s 701.82 MiB/s 722.01 MiB/s]
                     change:
                            time:   [+3.5656% +6.4151% +9.4428%] (p = 0.00 < 0.05)
                            thrpt:  [-8.6281% -6.0284% -3.4428%]
                            Performance has regressed.
    Found 1 outliers among 100 measurements (1.00%)
      1 (1.00%) high severe
    prost/google_message1_proto3/json_decode
                            time:   [1.6379 µs 1.6476 µs 1.6602 µs]
                            thrpt:  [236.67 MiB/s 238.48 MiB/s 239.89 MiB/s]
                     change:
                            time:   [+0.9807% +1.5064% +2.0611%] (p = 0.00 < 0.05)
                            thrpt:  [-2.0195% -1.4840% -0.9711%]
                            Change within noise threshold.
    Found 8 outliers among 100 measurements (8.00%)
      4 (4.00%) high mild
      4 (4.00%) high severe
    
  2. iainmcgin commented on Apr 22, 2026

    @iainmcgin
    Collaborator

    @brancz "zero copy" in the README is shorthand for the payload bytes of length-delimited fields are not copied from the input buffer into owned Rust types during decode. Varints, tags, and the fixed-width scalars are read, not copied in the memcpy sense. If that distinction would be clearer in the README we can tighten the wording — the claim we're making is narrower than "nothing is read from the input," and we don't want it to read as if it were.

    What a view does avoid

    A view is a borrow over the input buffer. For every length-delimited field, a view stores a slice into the original bytes:

    • string field → &'a str (UTF-8 validated once, no allocation)
    • bytes field → &'a [u8]
    • Nested message → SubMessageView<'a> (another borrow, no recursion into owned types)
    • repeated T (packed or otherwise) → RepeatedView<'a, T> — an iterator-style wrapper over a slice of the input; elements are lazily materialized on .iter()
    • map<K, V> → MapView<'a, K, V> — same idea; the map is walked, not built

    The result is that decode_view on a typical message does zero heap allocations, makes one linear pass over the wire format to record per-field slice offsets, and defers all per-element decoding to the caller. The MessageView trait is the contract.

    What a view cannot avoid

    Protobuf's wire format is not a memory layout, so some re-encodings are inevitable:

    1. Varint + tag reads. Every field needs its tag and (for length-delimited fields) its length decoded to find the next field. That is a per-field constant cost independent of field size — typically 1 or 2 bytes of varint. buffa's view decoder does exactly this linear scan once, records a small per-field descriptor (wire offset + length for length-delimited types; or the decoded value for fixed-width scalars), and stops. It does not copy payload bytes.
    2. Scalars that are not length-delimited (int32, fixed64, bool, etc.) are decoded eagerly during the scan, because there's nothing meaningful to borrow — they are already their target Rust type. For fixed32 / fixed64 / sfixed* / float / double, this is a small-constant byte swap on big-endian hosts and a pointer-read on little-endian; there is no "zero copy" representation for these that would also be correct cross-platform.
    3. UTF-8 validation for string fields. A view holds &'a str, not &'a [u8], so we check once that the bytes are well-formed UTF-8. That is required by the protobuf spec and by Rust's str invariant — no encoding lets us skip it without giving up one of the two.
    4. Unknown-field scanning. Views preserve unknown fields the same way owned messages do; the skipping pass reads tag+length for each unknown field. Again, linear in field count, not in payload size.

    The scan itself is eager (not lazy), because the per-field offset table is what makes subsequent field access O(1) without rescanning. A fully lazy design would push that cost onto the first field access instead, which is a different trade-off; buffa does not take it.

    Re: prost's bytes feature: Fair point @dezyh — we should benchmark prost with its bytes::Bytes feature enabled, both for completeness and because it's the closest prost analogue to what a view does for bytes fields. We've now done that in PR #61. The new prost (bytes) benchmark variant applies prost-build::Config::bytes(["."]) (substituting bytes::Bytes for every bytes field) and decodes from bytes::Bytes input, which is what actually exercises prost's zero-copy copy_to_bytes slicing path. The benchmark suite also now includes a deliberately bytes-heavy MediaFrame message (primary bytes body + repeated bytes chunks + map<string, bytes> attachments) so the feature has something to work with.

    Raw throughput (MiB/s, higher is better; Intel Xeon Platinum 8488C, task bench-cross in Docker; prost (bytes) deltas are relative to prost, buffa (view) deltas relative to buffa owned):

    Message prost prost (bytes) buffa owned buffa (view)
    ApiResponse 756 676 (−11%) 862 1,475 (+71%)
    LogRecord 712 676 (−5%) 722 1,984 (+175%)
    AnalyticsEvent 254 194 (−24%) 199 320 (+61%)
    GoogleMessage1 956 931 (−3%) 1,014 1,341 (+32%)
    MediaFrame 9,648 23,516 (+144%) 16,816 73,004 (+334%)

    Observations:

    1. prost (bytes) tracks default prost within noise on the larger messages and is ~10% slower on ApiResponse. The slow-down on ApiResponse is real: the proto message is small (~128 bytes/payload) and per-message fixed cost dominates, and enabling bytes results in extra work for each message decode. Specifically, the atomic fetch_add/fetch_sub around each payload.clone() plus the BytesAdapter-generic codepath run even if there are no bytes fields. This adds up to nontrivial overhead without any zero-copy payoff to offset it. On larger single messages that overhead is a smaller fraction of the total decode cost (which explains the results for LogRecord/GoogleMessage1), and with deeply nested small messages like AnalyticsEvent, the overheads make bytes significantly worse (~24% slower here).

    2. On MediaFrame, prost (bytes) is ~2.4× faster than default prost. This is exactly the payoff the feature advertises — the primary body, the chunks, and the attachments are all sliced from the input buffer rather than copied into fresh Vec<u8>s. This overcomes the overheads seen for the other message types that don't contain bytes fields.

    3. buffa views remain ~3× faster than prost (bytes) on MediaFrame. To explain where the remaining gap comes from, we ran perf stat on a native --measurement-time 6 run of each variant on MediaFrame decode. Sample results:

      variant GiB/s cyc/msg ins/msg L1-miss/msg IPC
      prost 9.1 3,929 10,470 113 2.67
      buffa owned 16.6 2,101 5,679 91 2.70
      prost (bytes) 21.9 1,632 4,625 5 2.83
      buffa (view) 69.8 508 1,675 4 3.30

      prost (bytes) closed the cache-miss gap — L1-dcache misses drop 113 → 5, which is the allocator traffic from cloning bytes payloads disappearing. But it still runs 2.8× more instructions per message than views: ~3k instructions per message is spent on String allocation for frame_id and content_type, UTF-8 validation of those strings, the HashMap<String, Bytes> the attachments map still builds (with hashed + heap-copied keys), and the Vec<Bytes> the chunks field still collects. Views skip all of it: the two strings stay &'a str, attachments is a MapView<'a, &'a str, &'a [u8]> that walks the input lazily on iter, and chunks is a RepeatedView<'a, &'a [u8]>.

    So prost (bytes) and buffa views are solving different (overlapping) problems: .bytes(["."]) is the correct prost feature for "don't copy bytes payloads", and it delivers cleanly on that. Views go further by not building the owning Rust collections for strings, maps, and repeateds either. When those fields aren't in your schema, the two should converge; when they are, the gap is what we see above.

  3. brancz commented on Apr 24, 2026

    @brancz
    Author

    I think this could be even clearer. It’s starting to make sense though, the point is that the input buffer never gets copied (hence zero copy), but it’s not zero-heap-allocations since the offset mask still needs to be computed.

  4. locked and limited conversation to collaborators on Apr 24, 2026
  5. converted this issue into a discussion #63 on Apr 24, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions