Blob Transport
What a blob is
Section titled “What a blob is”A blob is an opaque, content-addressed byte string. It is identified by a
BlobId — a 32-byte hash derived from the blob’s content — and it is displayed
as a hex string. Because the identifier is a hash of the bytes, a BlobId
both names a blob and lets any holder verify that the bytes it received are the
bytes that were asked for.
Blobs are the protocol’s mechanism for moving payloads that are too large, too binary, or too rarely-read to belong inside replicated state: WASM application bundles, file attachments, media, and similar. The state tree carries the small, hot, causally-ordered data; blobs carry the bulk, referenced by id.
pub struct BlobId(Hash); // 32-byte digest, hex on displayWhere blobs sit relative to state
Section titled “Where blobs sit relative to state”Blobs are not folded into a scope’s Merkle root. The convergence machinery
described in Storage and Sync compares
scope_root over state entities; blob bytes never enter that hash. State
references a blob only by storing its BlobId as a value. Two consequences
follow:
- Announcing, transferring, or garbage-collecting a blob does not change any scope root and cannot, by itself, cause or repair state divergence.
- A node can hold a
BlobIdin state without holding the bytes. The bytes are fetched lazily, on demand, over the transfer paths below — and a peer that serves them is checked against the referencing context, not against the state root.
Blob bytes live in their own storage, separate from the replicated entity tree:
metadata in the Blobs column and the raw chunk files on disk.
// crates/store/src/types/blobs.rs — value stored in the `Blobs` columnpub struct BlobMeta { pub size: u64, // total size in bytes pub content_hash: ContentHash, // SHA-256 over the whole file content pub links: Box<[key::BlobMeta]>, // chunk blobs, in order (empty for a leaf chunk)}Chunked storage layout
Section titled “Chunked storage layout”On write, the blob store reads the input stream in fixed CHUNK_SIZE = 1 MiB
pieces. Each chunk is persisted as its own blob whose BlobId is the hash of
that chunk’s bytes. The root blob is then the ordered list of those chunk
ids: its BlobId is the hash of the concatenated chunk ids, and its BlobMeta
carries the total size, the whole-content content_hash, and the links to the
chunk blobs. A blob that fits in a single chunk still gets a root whose id is
derived from its one chunk id.
root BlobId = hash( chunk_id[0] ‖ chunk_id[1] ‖ … ) // BlobMeta.linkschunk BlobId = hash( chunk_bytes ) // 1 MiB each, last may be shortBlobMeta.content_hash = hash( whole file content ) // distinct from BlobIdBecause a chunk’s id is the hash of its bytes, two blobs that share identical
1 MiB chunks resolve to the same chunk BlobId and are stored once —
deduplication falls out of content addressing with no extra bookkeeping. A
delete(blob_id) path exists to drop a blob’s stored bytes
(crates/store/blobs/src/lib.rs), but content-addressed garbage collection of
chunks no longer referenced from state is not yet implemented: removing a state
reference does not by itself reclaim the bytes.
The guest blob API
Section titled “The guest blob API”Applications never touch storage or the network directly. They use host
functions exposed to the WASM guest. Writing is a streaming
create → write* → close sequence that yields a BlobId; reading is
open → read*. A blob can be published to a context for discovery with
announce_to_context.
blob_create() -> fdblob_write(fd, src: Buffer) -> bytes_writtenblob_close(fd, dst: BufferMut[32]) -> 1 // writes the final BlobId into dst[0..32]blob_open(blob_id: Buffer[32]) -> fdblob_open_in_context(blob_id: Buffer[32], context_id: Buffer[32]) -> fd // 0 if unavailableblob_read(fd, dst: BufferMut) -> bytes_read // 0 at end of blobblob_announce_to_context(blob_id: Buffer[32], context_id: Buffer[32]) -> 1-
blob_create()opens a write handle and returns an integer file descriptor (fd). Internally this spins up a background task that consumes written chunks and feeds them into the blob store. -
blob_write(fd, src)appends one chunk. The number of bytes written equals the source buffer length. Each call’s buffer must not exceedmax_blob_chunk_size. -
blob_close(fd, dst)finalises the upload, computes theBlobId, and writes those 32 bytes into the guest bufferdst. The handle is consumed. For a read handle,closesimply releases it. -
blob_open(blob_id)opens a read handle for an existing blob and returns anfd. The bytes are streamed lazily — opening does not yet fetch anything. Local-only: returns0if this node does not already hold the blob. -
blob_open_in_context(blob_id, context_id)isblob_open’s network-aware sibling: if the blob is not already local, it consults the context’s peers before giving up, and returns anfdfor the same streaming read once bytes are available. It returns0only when the blob is available neither locally nor from any peer. -
blob_read(fd, dst)copies the next bytes intodstand returns the count, which may be less than the buffer (a partial chunk is buffered internally and carried to the next call). A return of0means end of blob. The read position only advances after the copy into guest memory succeeds, so a failed copy can be retried without losing data.
Limits and errors
Section titled “Limits and errors”| Limit | Default | Enforced by |
|---|---|---|
max_blob_handles |
100 |
blob_create / blob_open (open handles per execution) |
max_blob_chunk_size |
10 MiB |
blob_write (per-write) and blob_read (per-read buffer) |
The host functions surface these errors:
BlobsNotSupported— the node is configured without a blob/network client (the API is unavailable, e.g. in a bare VM).TooManyBlobHandles— opening would exceedmax_blob_handles.InvalidBlobHandle— unknownfd, or using a write handle where a read handle is required (or vice versa).BlobWriteTooLarge/BlobBufferTooLarge— the write chunk or the read buffer exceedsmax_blob_chunk_size.InvalidMemoryAccess— a descriptor buffer is out of bounds, or theBlobIddestination forblob_closeis not exactly 32 bytes.
Announce and discover
Section titled “Announce and discover”A blob is announced so the context’s availability nodes can prefetch the
bytes. blob_announce_to_context(blob_id, context_id) looks up the blob’s size
from local metadata and sends a one-message notice — { blob_id, context_id, size } — over a dedicated stream protocol,
/calimero/blob-announce/1.0.0 (Networking), to each
of the context’s TEE members (ReadOnlyTee or RelayTee) and to nobody else.
Direct streams to a chosen set, never a topic publish: gossipsub flood_publish
fans every publish to every subscriber, so broadcasting would tell the whole
context about every blob. And the announce is the ONLY place the blob→context
association exists — BlobMeta is keyed by blob id alone, and state deltas
carry opaque values — which is why prefetch hangs off it rather than off an
enumeration of the context’s state.
A receiver prefetches only if it is itself a TEE member of that context — directly, or by inheritance from any ancestor group up to the namespace root, which is how a root-admitted fleet node covers the contexts of the subgroups it follows; read from its own governance state, never taken from the announcement — does not already hold the blob, and the advertised size is within the 500 MiB transfer cap. At most 2 prefetches run at once, and announcements arriving while both slots are busy are dropped rather than queued — a dropped prefetch costs availability for one blob, never correctness.
Announcing is best-effort in the same spirit, and the producer does not wait
for it: the notices are sent detached, so blob_announce_to_context returns
once the work is scheduled, not once it is delivered. A refused, failed or
undelivered announce never fails — nor delays — the write that produced the
blob. Anything that needs to know an availability node actually holds the bytes
must observe that on the availability node.
Known gap: an availability node that is offline when a blob is announced never learns about it and has no catch-up path. The blob stays findable by probing its original holder, so correctness is unaffected, but availability for that blob degrades until something announces it again.
Discovery is a separate mechanism, and does not read any announcement. A node
that holds a BlobId in state but not the bytes asks the context’s peers
directly: it takes the subscribers of
the context’s gossipsub topic as candidates and sends each one a probe — a
BlobRequest whose response header is read and whose chunks never are, so the
peer answers found without transferring anything.
Probes go out in batches of 8, concurrently within a batch and sequentially
between batches, stopping at the first peer that answers found: true. The
winner is then asked for the bytes over the transfer protocol below.
Candidates are ordered before they are batched: availability nodes first, then peers that recently served this node a blob in the same context, then the rest in whatever order the subscriber set gave them. On a context large enough that the sweep runs out of time or hits its ceiling (below), that ordering decides which peers are asked at all, not merely in what order: an availability node listed last by the subscriber set might never be probed, while first it answers in one round trip.
The recent-provider tier is a hint, not state. It is memory-only — nothing
is persisted and nothing is replicated — and bounded both per context and in how
many contexts it tracks, so it cannot grow with network activity. A node that
has just started has an empty one and orders candidates exactly as it would
without it. It records a successful fetch, not a probe that answered found: true: a probe says a peer claims custody, whereas a completed fetch says it
served bytes that hashed to the id asked for, which is the only one of the two a
lying or unreachable peer cannot manufacture.
The search is bounded three ways, and all three are load-bearing:
| Bound | Value | What it stops |
|---|---|---|
| Width | 8 probes outstanding | Fanning a miss out to the whole subscriber set at once |
| Duration | 30 s over the whole path | A search outliving the HTTP request that asked for it |
| Ceiling | 256 candidates per sweep | An instantly-answering 10 000-peer context turning one miss into 10 000 probes |
The duration is the bound that normally ends a sweep. Batches run until the 30 s budget is gone or the candidates are exhausted; how far that gets depends entirely on how fast peers answer. On a healthy network a probe answers in a few milliseconds, so the budget comfortably walks a whole context; against peers that stall to the 5 s probe timeout it manages about six batches.
The ceiling is a safety valve, not the normal stop. It is also where the deliberate incompleteness now lives: a blob held only by a peer beyond candidate 256 still reads as absent. That remains the escalation trigger for this design — a context that big needs real provider state (an anchor, a rendezvous, a sharded index), not a longer linear scan.
The cost of that incompleteness is asymmetric on purpose. A miss on a large context can now cost up to 256 probes instead of 32, but a miss is latency and the deadline already caps it, whereas a wrong “nobody has it” is a 404 for a blob that exists on a peer the node simply never asked.
A winner that then fails to deliver — a dropped connection, a timeout, or bytes whose recomputed id does not match — is not evidence that the blob is gone, so the search resumes at the next candidate instead of reporting absence. And an empty candidate list is not evidence either: a node that has just joined may simply not have been grafted onto the context topic yet. An empty sweep is therefore retried up to 6 times, re-resolving the subscriber set each time, with the gap starting at 100 ms and doubling to a 2 s cap. A sweep that had candidates and still found nothing is not repeated — re-asking peers that already said no is repeating work, not making progress.
Asking is authoritative in a way a record is not: custody is a property of a
peer’s own blob store, so the answer cannot be stale, cannot expire, and is not
limited to whichever single peer happened to announce. Any node that has the
bytes — including one that only ever downloaded them — is discoverable, and a
holder that restarts stays discoverable. A probe deliberately does not
distinguish “not authorised” from “not held”: both answer found: false, which
costs a wasted candidate and never a wrong fetch.
The Kademlia record, and why it is still written
Section titled “The Kademlia record, and why it is still written”The Kademlia provider-record API (announce_blob_to_kad / find_blob_providers)
is deprecated, and nothing in the tree reads it — discovery is probing, and
only probing.
The write survives anyway, and announce_blob_to_network performs it
alongside the availability-node notice. A peer still running the pre-probe code
discovers blobs by DHT lookup and by nothing else, so an upgraded node that
stopped writing the record would become invisible to it. The reverse direction
needs no such help: an un-upgraded peer answers an ordinary blob request, so a
probing node still finds it. Only the write is load-bearing for mixed versions,
which is why only the write was kept. Both routes sit behind the same
withholding gate, so an http-mode node advertises application bytecode by
neither.
The record is opportunistic in a way probing is not — only some paths publish one, and a restart drops what a node held — which is exactly why nothing reads it any more. It goes once no un-upgraded peer remains.
Point-to-point transfer over the dedicated stream
Section titled “Point-to-point transfer over the dedicated stream”Once a provider peer is known, the bytes move over a dedicated libp2p stream protocol, separate from the sync stream:
pub const CALIMERO_BLOB_PROTOCOL: StreamProtocol = StreamProtocol::new("/calimero/blob/0.0.3");pub const MAX_MESSAGE_SIZE: usize = 8 * 1_024 * 1_024; // 8 MiB framed-message capThe requester opens a stream on CALIMERO_BLOB_PROTOCOL, sends a single
BlobRequest, and reads a BlobResponse header followed by a sequence of
BlobChunk frames terminated by an empty chunk.
struct BlobRequest { blob_id: BlobId, context_id: ContextId, auth: Option<BlobAuth> } // JSONstruct BlobResponse { found: bool, size: Option<u64> } // JSONstruct BlobChunk { data: Vec<u8> } // BorshThe provider streams its stored chunks straight from the blob store, one
BlobChunk per stored chunk, then sends an empty BlobChunk to mark the end.
If the provider does not hold the blob it replies BlobResponse { found: false }
and sends no chunks.
Authorization
Section titled “Authorization”Blob bytes are not served to anyone who asks.
Application bytecode has a gate ahead of authorization: a node whose
[registry] mode is http resolves applications from its
own registry and is therefore not a source of their bytes, so it neither announces nor
serves them and the request reads as “not held” (NodeClient::may_share_blob). Only
dht-mode nodes serve application bytecode. User-data blobs are unaffected in both
modes, and everything below applies to them either way.
For the blobs a node will serve, the provider authorizes each BlobRequest before
streaming:
- Public bundles. If the requested
blob_idis the context’s application bytecode or compiled artifact (per the authoritative context config), access is granted with no signature. This is what lets a joining node fetch the code it needs before it has an identity in the context. - Signed member requests. Otherwise the request must carry a
BlobAuth: the requester’spublic_key, a Unixtimestamp, and an Ed25519signatureover theBlobAuthPayload { blob_id, context_id, timestamp }. The provider checks the signature, checks that the timestamp falls inside the replay window (30 s in the past, 10 s in the future), and checks that the public key is a current member of the context. Any failure denies access.
struct BlobAuth { public_key: PublicKey, signature: [u8; 64], timestamp: u64 }struct BlobAuthPayload { blob_id: [u8; 32], context_id: [u8; 32], timestamp: u64 }Timeouts and framing
Section titled “Timeouts and framing”Both ends bound the transfer so a stalled peer cannot pin resources:
| Side | Bound | Value |
|---|---|---|
| Requester | whole transfer | 60 s |
| Requester | per chunk received | 30 s |
| Provider | whole serve | 300 s |
Every frame is length-delimited and capped at MAX_MESSAGE_SIZE (8 MiB) by the
stream codec. Since stored chunks are 1 MiB, a BlobChunk frame stays well
under that cap.
Worked example: a 5 MB image
Section titled “Worked example: a 5 MB image”To see the pieces fit together, follow a 5 MB image from the app that stores it to a peer that fetches it.
-
Store it. The producing app calls
blob_create(), streams the image in with repeatedblob_write(fd, …), and callsblob_close(fd, dst). The blob store re-chunks the byte stream into roughly five 1 MiB storage chunks (the last one short), andclosewrites back the rootBlobId— the hash of the ordered chunk-id list. The app keeps that id (typically by storing it in state). -
Announce it.
blob_announce_to_context(blob_id, context_id)looks up the blob’s size locally and sends{ blob_id, context_id, size }over/calimero/blob-announce/1.0.0to each of the context’s availability nodes. Each one that is online, does not already hold the image, and has a free prefetch slot fetches it, so the image ends up on an always-on node as well as the producer. -
Discover it. A second node holds the
BlobIdin state (it synced the reference) but not the bytes. It takes the context topic’s subscribers as candidates and probes them 8 at a time — aBlobRequestwhosefoundanswer it reads and whose chunks it does not — until one says yes. Any node holding the bytes qualifies, not just the announcer. -
Fetch it. The fetcher opens
/calimero/blob/0.0.3to that peer and sends aBlobRequest { blob_id, context_id, auth? }. The provider authorizes the request, repliesBlobResponse { found: true, size }, then streams its stored chunks asBlobChunkframes — one per 1 MiB chunk — and a final emptyBlobChunkto mark the end. -
Verify it. The fetcher concatenates the chunks, recomputes the
BlobId, and accepts the bytes only if it matches the id it asked for. It stores the image locally and can now serve (and announce) it itself.
Blob transfer during sync
Section titled “Blob transfer during sync”Blobs can also move inside a sync session over the sync stream
(/calimero/stream/0.0.3), rather than the dedicated blob protocol. This path
is used when a sync needs a blob it does not yet hold — for example staging an
application upgrade. It rides the same encrypted, sequenced envelope as every
other sync exchange (Sync), so the wire messages differ from
the dedicated protocol:
InitPayload::BlobShare { blob_id: BlobId } // request (and the responder's ack)MessagePayload::BlobShare { chunk: Cow<'a, [u8]> } // each chunk; empty chunk ends the streamThe initiator sends an Init carrying BlobShare { blob_id }; the responder
acks with the same blob_id, then streams Message frames each carrying a
BlobShare chunk, ending with an empty chunk. Unlike the dedicated protocol,
these frames are encrypted under the per-session shared key and ordered by the
sync sequencer, and the chunk payload is raw bytes (not a Borsh BlobChunk). A
responder that does not hold the blob replies with an opaque error so the
initiator fails fast and retries after its backoff rather than blocking.
Summary
Section titled “Summary”- A blob is content-addressed immutable bytes named by
BlobId, stored outside the state tree in theBlobscolumn plus on-disk chunk files, and re-chunked to 1 MiB internally. - Blobs are referenced from state by id only; they are never folded into
scope_root, so blob movement neither creates nor repairs state divergence. - Guests use
create → write → close(yields aBlobId),open → read, andannounce_to_context, bounded bymax_blob_handlesandmax_blob_chunk_size. - Discovery is a bounded probe of the context’s peers, availability nodes first;
producers announce new blobs to those availability nodes over
/calimero/blob-announce/1.0.0so they prefetch. Transfer runs either over the dedicated/calimero/blob/0.0.3stream (BlobRequest→BlobResponse→BlobChunk*) or, during sync, as encryptedBlobShareframes — both gated by context-membership authorization and both self-verifying by content address.