@@ -2249,3 +2249,91 @@ This module is experimental and subject to change.
22492249```{eval-rst}
22502250.. py:module:: torch.distributed.checkpoint.state_dict
22512251```
2252+
2253+ ```{eval-rst}
2254+
2255+ One-sided tensor transports (experimental)
2256+ --------------------------------------------------
2257+
2258+ ``torch.distributed._transport`` moves registered tensor byte ranges between
2259+ independently managed workers. It does not require a process group, global rank
2260+ assignment, or matching receives.
2261+
2262+ An endpoint connects to one peer. Exchange ``bind()`` bytes and remote memory
2263+ descriptors through a trusted application control plane. Exchange descriptors only with authorized peers. A write copies from a local view to the remote base; a read
2264+ copies from the remote base to a writable local view. Offsets and lengths are bytes;
2265+ shape and dtype agreement is the application's responsibility.
2266+
2267+ Synchronous reads/writes return zero; ``async_op=True`` returns
2268+ :class:`torch.distributed.Work`. Successful completion means the transfer completed,
2269+ not that the remote application consumed or acknowledged the data. Asyncio callers
2270+ can use ``read_async``, ``write_async``, or ``wait_all``. Registration remains valid
2271+ until unregistration or close, and tensors must not be resized or have their storage replaced.
2272+
2273+ CUDA stream semantics, graph capture, tracing, batching, remote slicing, and
2274+ rank-based bootstrap helpers are outside this initial API. Rank-to-endpoint
2275+ lookup belongs in a separate control-plane adapter. Descriptor classes define explicit ``serialize()``/``deserialize()`` methods.
2276+ The built-in backends declare their fields in a versioned JSON envelope; binary
2277+ metadata is base64-encoded. Unknown fields, versions, backends, and invalid field
2278+ types are rejected. Tensor contents and native handles are never serialized.
2279+
2280+ .. autofunction:: torch.distributed._transport.new_transport
2281+ .. autoclass:: torch.distributed._transport.Transport
2282+ :members:
2283+ .. autoclass:: torch.distributed._transport.Memory
2284+ :members:
2285+ .. autoclass:: torch.distributed._transport.MemoryView
2286+ :members:
2287+ .. autoclass:: torch.distributed._transport.MutableMemoryView
2288+ .. autoclass:: torch.distributed._transport.RemoteBuffer
2289+ .. autofunction:: torch.distributed._transport.wait_all
2290+ .. autofunction:: torch.distributed._transport.available_transports
2291+ .. autofunction:: torch.distributed._transport.register_transport
2292+
2293+ NIXL backend
2294+ ~~~~~~~~~~~~
2295+
2296+ The initial backend is NIXL, installed separately with ``pip install nixl``.
2297+ Its default UCX plugin supports CPU and CUDA memory, subject to the installed
2298+ NIXL/UCX build and hardware.
2299+
2300+ ``unregister_memory(memory)`` deregisters a local allocation without closing the
2301+ transport. Wait for local transfers first and coordinate with peers to stop remote
2302+ access; the backend cannot detect incoming DMA. All handles sharing a registration,
2303+ including previously created views and exported descriptors, become invalid.
2304+ Register again and exchange fresh descriptors before resuming transfers.
2305+
2306+ NIXL transfers return Work objects that retain their request handles until
2307+ completion. ``wait_all``, ``read_async``, and ``write_async`` await Work futures.
2308+ The NIXL adapter resolves those futures by checking native transfer status.
2309+ Each live transfer owns a distinct request handle.
2310+ Independent requests may overlap; explicitly wait before issuing dependent or
2311+ overlapping reads/writes. Completion ordering is not implicit.
2312+
2313+ NIXL's default wait timeout is 30 seconds. ``timeout`` is in seconds; ``None``
2314+ selects the backend default and zero polls without waiting. ``Work.wait`` instead
2315+ takes a ``datetime.timedelta``; its zero default selects the transfer's timeout.
2316+ Timeout and asyncio cancellation stop waiting, not DMA. The transport retains
2317+ pending requests and buffers, even if the caller drops its Work. Wait again or
2318+ successfully close before reusing buffers. Coordinate with peers before closing
2319+ exposed memory; close only drains locally submitted operations.
2320+
2321+ ``NIXLTransport.close_async`` awaits pending transfers before native cleanup.
2322+ A timed-out or cancelled close rejects new work and retains resources; retry
2323+ close to finish cleanup. Registrations keep the transport alive even after its last outgoing transfer,
2324+ because peers may still access exposed memory. Forgotten registrations may retain
2325+ resources indefinitely:
2326+ call ``unregister_memory`` or ``close`` after coordinating with peers.
2327+ Checking ``is_completed`` also releases completed requests.
2328+ For pending NIXL work, ``get_future`` requires a running asyncio loop; the loop
2329+ must remain running to drive that future. Completed work needs no event loop.
2330+
2331+ Native metadata, registration, request submission, status checks, and cleanup
2332+ calls execute synchronously on the calling thread. Their timeouts bound lock
2333+ acquisition and transfer completion waits, not execution inside NIXL. Python
2334+ cannot interrupt a blocked native call, even when it releases the GIL.
2335+
2336+ .. autoclass:: torch.distributed._transport.nixl.NIXLTransport
2337+ :members: close_async
2338+
2339+ ```
0 commit comments