Skip to content
33 changes: 33 additions & 0 deletions docs/serialization.md
Original file line number Diff line number Diff line change
Expand Up @@ -22,6 +22,8 @@ user = User.from_binary(data)

### Options

#### Unknown Fields

When a message is parsed from binary data containing field numbers it doesn't recognize, the unknown fields are stored internally and re-emitted during serialization.
This means a message can pass through an intermediary that doesn't know about newer fields without losing data.

Expand All @@ -37,6 +39,37 @@ data = user.to_binary(write_unknown_fields=False)
user = User.from_binary(data, ignore_unknown_fields=True)
```

#### Allocation Limit

Because protobuf serialization can create very compact binary payloads, it is possible for the memory usage of a
parsed message to differ drastically from the input number of bytes. When parsing untrusted payloads,
such as in an external-facing API server, this can allow malicious users to send small messages that take
a large amount of memory or potentially crashing the server. This is most pronounced in schemas with
repeated fields of message type with a large number of fields.

You can mitigate this using the `allocation_limit` option in `from_binary`. When set, an estimate of the
memory usage of a message is maintained while it is parsed, and if it goes over the limit, the parse fails
immediately before processing the entire payload. The limit is based on the schema and content, not the environment,
i.e., it charges a value for a string field based on the number of characters with a fixed overhead for a Python
string. The overhead can change between Python versions, so the allocation budget should not be considered
a precise value. Such changes are generally relatively small and fixed so an allocation limit determined for
one environment should generally work fine when i.e., updating Python.

The allocation budget is an upper bound, so it is possible that a message that would be under the limit is
rejected. For example, when parsing binary the string is charged with the number of utf8 bytes in the payload.
For ASCII strings, this will be the precise number, but for i.e. certain CJK characters, it overcharges by
~30%.

It is recommended to set an allocation limit for applications parsing untrusted payloads based on your target
memory usage. You may need to experiment with values to see the effective memory utilization due to potential
overcharging.

```python
user = User.from_binary(
data, allocation_limit=8 * 1024 * 1024
) # Roughly cap memory usage of parsed message to 8MB
```

## JSON

```python
Expand Down
56 changes: 56 additions & 0 deletions packages/protobuf-py-ext/src/budget.rs
Original file line number Diff line number Diff line change
@@ -0,0 +1,56 @@
use pyo3::{PyResult, exceptions::PyValueError};

// Approximate sizes of CPython heap allocations, measured with
// `sys.getsizeof` on 64-bit CPython 3.14. The budget guards against
// unbounded allocation from malicious payloads rather than providing exact
// accounting, so small inaccuracies across versions and builds are fine.

/// GC header allocated in front of every GC-tracked object.
pub(crate) const GC_HEAD_SIZE: usize = 16;
/// A float object.
pub(crate) const FLOAT_SIZE: usize = 24;
/// A 64-bit int object. Smaller ints are slightly smaller.
pub(crate) const INT_SIZE: usize = 36;
/// Header of a compact ASCII str. The UTF-8 byte length is charged on top
/// as an approximation of the payload.
pub(crate) const STR_OVERHEAD: usize = 41;
/// Header of a bytes object.
pub(crate) const BYTES_OVERHEAD: usize = 33;
/// An empty list.
pub(crate) const EMPTY_LIST_SIZE: usize = 56;
/// An empty dict.
pub(crate) const EMPTY_DICT_SIZE: usize = 64;
/// One appended list element: an 8-byte pointer slot.
pub(crate) const LIST_SLOT_SIZE: usize = 8;
/// One inserted dict entry: hash + key + value words plus growth slack.
pub(crate) const DICT_ENTRY_SIZE: usize = 40;
/// A `Oneof` wrapper object: `PyObject` header plus two object pointers.
pub(crate) const ONEOF_SIZE: usize = 32;

/// Tracks the approximate bytes of Python objects allocated while parsing a
/// message, raising an error once a configured limit is exceeded.
pub(crate) struct Budget {
current: usize,
max: usize,
}

impl Budget {
pub(crate) fn new(limit: Option<usize>) -> Self {
Self {
current: 0,
max: limit.unwrap_or(usize::MAX),
}
}

pub(crate) fn charge(&mut self, amount: usize) -> PyResult<()> {
self.current = self.current.saturating_add(amount);
if self.current > self.max {
Err(PyValueError::new_err(format!(
"allocation budget exceeded: needed {} bytes, limit is {}",
self.current, self.max
)))
} else {
Ok(())
}
}
}
3 changes: 3 additions & 0 deletions packages/protobuf-py-ext/src/constants.rs
Original file line number Diff line number Diff line change
Expand Up @@ -93,6 +93,8 @@ pub(crate) struct ConstantsInner {

/// The string `__new__`.
pub(crate) dunder_new: Py<PyString>,
/// The string `__basicsize__`.
pub(crate) dunder_basicsize: Py<PyString>,

/// Python types.
pub(crate) types: Types,
Expand Down Expand Up @@ -146,6 +148,7 @@ impl Constants {
values: PyString::new(py, "values").unbind(),

dunder_new: PyString::new(py, "__new__").unbind(),
dunder_basicsize: PyString::new(py, "__basicsize__").unbind(),

types: Types {
desc_field: mod_descriptors
Expand Down
Loading
Loading