Fleet HA Sidecar
The fleet HA sidecar is the Mero TEE half of hardware-attested fleet
admission. It is a Bash script
(mero-tee/ansible/roles/merotee/templates/fleet-sidecar.sh.j2) baked into the
ReadOnly TEE node image and run as a systemd service alongside merod. Its
job is to keep this node’s Calimero namespace membership in sync with what
the fleet control plane (MDMA) currently assigns to it: join the namespaces
MDMA wants it in, and leave the ones it withdraws.
The admission handshake itself — quote generation, TeeAttestationAnnounce,
MRTD-gated admission — lives in Calimero Core, not here. The sidecar is only
the trigger: it drives merod through meroctl and reports back to MDMA. See
Calimero Core → TEE Attestation & Fleet Admission
for the protocol it drives.
What it reads at startup
Section titled “What it reads at startup”Before it can poll, the sidecar has to know two things about itself, plus where to talk to MDMA. Configuration is baked at image build time via Ansible (so it is measured into the MRTD), not read from instance metadata; TLS to MDMA is always verified.
| Value | Source | Notes |
|---|---|---|
MDMA_URL |
Ansible build-time var (fleet_mdma_url) |
Sidecar exits if unset. |
FLEET_TOKEN |
Issued by MDMA on attested registration, stored at /mnt/data/fleet/fleet-token |
Never baked into the image. Sent as X-Fleet-Token once held; the two bootstrap routes take none. See How the node gets its fleet token. |
| PeerId | /mnt/data/calimero/default/config.toml → [identity].peer_id |
Read after merod is ready; retried once after a short sleep. |
| MRTD | /sys/class/misc/tdx_guest/measurements/mrtd:sha384 |
See below. |
| Executor account | meroctl --output-format json account show → accountId |
The account this node writes as. Read at startup and, if merod is not answering yet, retried on the loop every ACCOUNT_RETRY_INTERVAL (15s) until it does — see When the account is not readable at boot. See Delegated execution. |
| Relay URL | Instance metadata key relay-url |
The one value not baked at build time — it is per-node. See below. |
When the account is not readable at boot
Section titled “When the account is not readable at boot”The executor account’s value is fixed for the life of the node. Its availability at the moment it is first read is not, and conflating the two used to make a node permanently, silently useless as a relay.
The subtlety is worth stating, because the obvious explanation is the wrong
one: this is not merely After=merod.service ordering a start without
waiting for readiness. wait_for_merod runs first and does block until merod
answers. The gap is narrower — wait_for_merod polls meroctl peers, while
the account is read with meroctl account show. A merod that already serves
peers has not necessarily minted this node’s account yet, so one call can
succeed while the other still fails. When it did, the startup read left
EXECUTOR_ACCOUNT empty and nothing ever read it again.
Everything downstream hangs off that one value:
reconcile_registrationreturns immediately without an account — there is no identity to bind — so the node never registers.- No registration means MDMA creates no
NodeCertificaterow. - No certificate row means the dispatcher publishes no A record and orders
no certificate, so the node’s
*.relay.<zone>name answersNXDOMAIN. _record_relay_identityrefuses to record a relay URL for an unregistered peer, so the node is never advertised as a relay.
Throughout, the node replicates and polls perfectly happily. It looks healthy from every angle except the one that matters, which is why this is worth stating explicitly rather than leaving to the code.
reconcile_executor_account closes it: while the account is still unreadable
it re-reads once every ACCOUNT_RETRY_INTERVAL (15s), logs
Fleet node executor account: … (readable after Ns) when it succeeds, and then
costs nothing at all for the life of the process. Covered by
scripts/ci/tests/fleet-sidecar-account-retry-test.sh.
And reconcile_registration now says when it is in that state. It opens by
returning early if it has no PeerId or no account, which was correct but silent:
at 1 Hz, a node could sit there for its whole life without a word. It is
edge-triggered, like the authorship re-report — one line when it starts
blocking, naming which field is missing and what that costs, one when it clears,
and nothing in between.
Waiting for merod, with a bound and a voice
Section titled “Waiting for merod, with a bound and a voice”wait_for_merod polls meroctl peers until it answers. It used to do so
forever, with no line inside the loop and stderr discarded, so a node whose
merod never came up stopped there permanently and said nothing — no poll, no
registration, no certificate, no DNS record, no relay — while looking healthy
from outside, because merod’s own port was answering and only meroctl was not.
On a profile with no shell there is then no way left to ask what happened.
Three properties now:
- Bounded. Past
MEROD_WAIT_TIMEOUT(600s) it gives up and the caller carries on. Carrying on is deliberate:get_peer_idreadsconfig.tomldirectly rather than through merod, so the sidecar can still register, still poll, and still be seen by MDMA, whilereconcile_executor_accountkeeps retrying the read that failed. Blocking here turns a slow merod, or one broken subcommand, into a node that is permanently invisible. - Progress. A line every
MEROD_WAIT_LOG_EVERY(30) attempts, so the journal distinguishes “still waiting, 90s in” from “died before it got here”. stderrkept.meroctl’s own complaint is the whole diagnosis — a missing home directory, a refused connection and an unparseable config all look identical once2>/dev/nullhas eaten them.
The observability credential comes from MDMA, not the image
Section titled “The observability credential comes from MDMA, not the image”No secret is baked into the node image, and the bearer token vector and vmagent present to the observability front end is the worked example of that rule.
The image is an artifact whose distribution is not fully controlled, so treat anything inside it as readable. What the image carries is the code to ask for a credential.
Secret Manager is not an alternative, and not for policy reasons: MDMA creates
these instances with no serviceAccounts field at all, so an MDMA node has
no GCP identity. Neither gcloud nor the metadata server’s token endpoint has
anything to authenticate as. That is what the provided provider in
fetch_secret.sh exists for — it reads a token already on disk rather than
fetching one.
Two delivery paths, and why there are two
Section titled “Two delivery paths, and why there are two”| path | when | written by | reaches |
|---|---|---|---|
observability-token instance metadata |
at creation | the MDMA dispatcher | calimero-init, before any unit starts |
should-join response |
on any poll | the MDMA manager | a node that is already running |
Metadata is the bootstrap. The poll is the refresh. Neither is sufficient alone.
The poll was once the only path, and that made the observability path depend on the thing it observes:
the poll needs a fleet token → the fleet token needs attested registration → registration needs the sidecar's main loop runningA node whose sidecar never reached its first poll could therefore never ship a
log line — which is exactly the node whose logs you need. calimero-init’s own
journal, where such a node says why, was unreachable for the same reason.
Reading the credential from instance metadata at boot breaks that circle
completely: vector comes up authenticated whether or not the sidecar ever runs,
ever registers, or can even reach MDMA.
The poll stays because instance metadata is written at creation and is never backfilled, so it is the only way a rotated token reaches a running node without recreating it.
The two agree by construction. Both write /etc/vector/provided_token with
printf '%s' and no trailing newline, and the node reports logs_token_fp —
the SHA-256 of the token it holds, never the token — so MDMA sends one only
while that disagrees with the configured value. A node booted with a current
token never has the credential ride the poll at all. At 1 Hz, re-sending
unconditionally would put it on the wire 86,400 times a day.
The trade-off
Section titled “The trade-off”Putting the token in instance metadata means one fleet-wide bearer token in a place any process on the box can read, rather than a per-node credential handed over an attested channel. That is weaker than delivery-by-attestation alone, and it was chosen deliberately.
What caps it:
- The standing rule holds. Instance metadata is not the image: it is per-instance, and it rotates without a release or an MRTD re-allowlist.
- The
vmauthuser the token names is scoped write-only — insert paths only, no read access to logs or metrics. A node that leaks it lets someone add to the store, not read the fleet’s telemetry. That matters precisely because a node is a machine nobody can log into to rotate a credential in place.
Against that: a node with no logs cannot be debugged at all, on a profile with no shell to debug it from.
How it is written to disk
Section titled “How it is written to disk”Both writers create the token inside ( umask 077; ... ), and neither relies on
a chmod afterwards.
That distinction is the whole point. chmod 600 after the write closes the
window; it does not prevent it. Neither calimero-init.service nor
fleet-sidecar.service sets UMask=, so systemd’s default 0022 applies and the
file is created 0644 — the credential sits world-readable on disk until the
chmod lands. The chmod is kept, but as a belt-and-braces assertion of the final
mode rather than as the thing that sets it.
It applies to both writers because they write the same file. A difference in
how /etc/vector/provided_token is created depending on which code path got
there first would be a difference nobody would think to look for.
Low severity on a locked image — there is no shell and no unprivileged user with a reason to be looking — so this is hygiene rather than an exploited hole. The cost of getting it right is one subshell.
Two guards, because one of them could not reach both cases.
fleet-sidecar-logs-token-test.sh shadows chmod to observe the mode as
created, which a stat of the final file cannot see — but it can only test
what it can call, and calimero-init is a linear boot script with no sourceable
functions. secret-file-umask-test.sh therefore checks both templates
statically. CSRs and certificates are deliberately excluded: they are public
material.
Rotation
Section titled “Rotation”Change MDMA_OBSERVABILITY_LOGS_TOKEN on both the manager and the dispatcher —
they must carry the same value, or the fingerprint never matches and the
credential rides every poll forever. Running nodes pick it up on their next poll
and both shippers are reconfigured together; nodes created afterwards get it at
boot.
Reconfiguring only vector would leave vmagent presenting the superseded token, and every sample would be rejected at the sink — silently, on a node with no shell to notice it from. The sidecar therefore always does both.
Covered by scripts/ci/tests/fleet-sidecar-logs-token-test.sh.
Reading a locked node’s sidecar log
Section titled “Reading a locked node’s sidecar log”There is no shell on a confidential VM, so /run/calimero/fleet-sidecar.log is
unreadable by anyone once an image is running — including whoever is trying to
work out why a node never became a relay.
log tees to stdout as well as that file, and the unit sets no
StandardOutput=, so every line also reaches journald under
_SYSTEMD_UNIT=fleet-sidecar.service. What decides whether it leaves the box is
include_units in
ansible/roles/merotee/files/vector-partial-merod.yaml, which Vector ships to
VictoriaLogs. fleet-sidecar.service is in that list; a unit left out of it is
a subsystem nobody can debug on a locked image, whatever it writes locally.
include_units carries calimero-init.service, merod.service,
mero-auth.service, traefik.service, fleet-sidecar.service,
vector.service and vmagent.service.
calimero-init.service is the one to know about. It starts every other unit in
that list, and it starts the sidecar with systemctl start --no-block, which
logs Queued non-blocking start whether the unit then came up or died on its
first line. A sidecar that never ran at all is invisible from anywhere else.
Two things have to be true before any of it leaves the node, and until they are the config above is inert:
logs-endpointinstance metadata is set — the dispatcher writes it only whenMDMA_TEE_METADATA_LOGS_ENDPOINTis configured, and metadata is set at creation only, so an existing node never picks it up.- The node holds the bearer token, per the section above — from
observability-tokenmetadata at boot, or from a poll.
The destination is the public victoria-lb front end, read back through
Grafana; the internal victoria-logs-int ALB resolves to private addresses in
the AWS VPC and is not reachable from a GCP node.
How the node gets its fleet token
Section titled “How the node gets its fleet token”No credential is baked into the image. A node boots holding nothing and earns its token by attesting.
boot (no secret) → GET /api/fleet/nodes/challenge unauthenticated → POST /api/fleet/nodes/register unauthenticated; carries a TDX quote ← { "fleet_token": "f1.<peer id>.<hmac>" } → every other /api/fleet/* call X-Fleet-Token: <that token>The token is stored at /mnt/data/fleet/fleet-token — on the encrypted data
disk, never the boot disk (older images kept it under /var/lib/calimero/;
calimero-init moves it on the first boot of a newer image) — mode 0600, and survives
restarts so a node does not re-register on every boot.
Why the bootstrap routes are open. They have to be — requiring a token to
obtain a token is circular. It costs nothing, because the shared secret that
used to sit in front of them was strictly weaker than the check behind them: a
challenge is a sealed nonce, worthless without a quote, and registration
verifies a TDX quote whose MRTD must be in published-mrtds.json. Every other
route still requires a token, and
scripts/ci/check-fleet-token-headers.py enforces that every function calling
/api/fleet/* offers the header.
What this replaced, and why. One shared secret for the whole fleet, baked in
at build time. It was never the security boundary — it proved “some fleet node”,
never which one; that is what attestation is for, and everything that mattered
(certificate issuance, relay identity, assignment confirmation) was already
gated on the quote. But it sat in an artifact whose distribution is not fully
controlled, and being measured into the MRTD meant rotating it was an image
release plus a republished published-mrtds.json.
Per-node tokens fix both ends:
| shared, baked | per-node, issued | |
|---|---|---|
| in the image | yes | no |
| scope | whole fleet | one peer id |
| one node compromised | fleet-wide access | that node only |
| rotation | rebuild every image, re-allowlist MRTD | rotate MDMA_FLEET_TOKEN_SIGNING_KEY |
On the MDMA side the value is MDMA_FLEET_TOKEN_SIGNING_KEY — note the MDMA_
prefix, which is what the manager’s settings actually read. It signs the tokens
it issues; rotating it invalidates every token at once, and nodes re-register on
their next cycle and are issued new ones, with no image involved.
Retry is on a clock. A node with no token registers regardless of whether
its identity changed, since registering is the only way to get one — but at most
once per REGISTRATION_RETRY_INTERVAL (30s). Each attempt costs a TDX quote and
spends a single-use challenge, so a node MDMA will not issue to (measurement not
yet published, or no signing key configured there) would otherwise generate a
quote a second, forever.
Discovering the PeerId
Section titled “Discovering the PeerId”get_peer_id() reads the node’s PeerId straight from its config file at
/mnt/data/calimero/default/config.toml, under [identity].peer_id (via
tomllib, falling back to a grep/sed parse). It does not shell out to
meroctl for this — an earlier version that tried meroctl peers list failed
because peers takes no arguments and returns only a count.
Discovering its own MRTD
Section titled “Discovering its own MRTD”get_mrtd() reads this TDX guest’s own MRTD (the SHA-384 measurement of the
initial VM image) directly from the kernel-exposed sysfs attribute:
/sys/class/misc/tdx_guest/measurements/mrtd:sha384This is the same 48-byte value the Intel Quoting Enclave writes into a TDX quote
(quote.mrtd()), but the kernel surfaces it read-only, so no quote generation
or configfs-tsm round-trip is required. The script hex-encodes it to 96
lowercase hex characters with no 0x prefix — exactly the shape MDMA’s MRTD
allowlist compares against. MRTD is computed once at startup because it is
fixed for the life of the immutable node image.
The reconcile loop
Section titled “The reconcile loop”After a one-time wait_for_merod() (which blocks until
meroctl --output-format json peers exits 0) and the PeerId/MRTD reads above,
the sidecar enters an infinite loop with a one-second poll interval
(POLL_INTERVAL=1). Each cycle:
- Poll MDMA.
POST {MDMA_URL}/api/fleet/should-joinwith body{"peer_id": "...", "mrtd": "..."}(curl -sf --max-time 10, plus theX-Fleet-Tokenheader if a token is configured). - Gate on poll success. The response must be HTTP 200 and parse as an
object with an
assignmentsarray. Any curl error, non-2xx, timeout, or unparseable/wrong-shape 200 body aborts the cycle before any state change — the loop sleeps andcontinues, leaving the confirmed set untouched (see the safety gate below). - Compute
desired. The sorted set ofgroup_ids inassignments— every namespace MDMA currently assigns this node to. - Join toward
desired. For each assigned namespace not already in the local confirmed set, runjoin_group(see below). The moment local admission succeeds the namespace is recorded infleet-admitted.json— before the confirm is attempted, and regardless of whether it succeeds. ThenPOST{MDMA_URL}/api/fleet/confirmwith{"peer_id": "...", "group_id": "..."}. A namespace is added tofleet-confirmed.jsononly when both local admission and the MDMA confirm succeed. - Leave dropped namespaces. Compute
to_leave = admitted − desiredandleave_groupeach one (see below). This runs before the prunes, so the diff is taken whileadmittedstill holds the dropped entries. - Prune both sets. Set
confirmed = confirmed ∩ desiredandadmitted = admitted ∩ desired, and persist them — so a future re-add is recognized as new, and a namespace just left is not left again every second. - Report the context inventory, over the pruned set — but on its own much slower clock than this loop, and immediately when the confirmed set itself changed. See The context inventory.
- Sleep one second and repeat.
The confirmed set is persisted at /mnt/data/fleet/fleet-confirmed.json
(a deliberately separate path from any pre-upgrade legacy state file, and a
different semantics: it records “admitted and confirmed”, not merely
“attempted”).
The admitted set lives beside it at /mnt/data/fleet/fleet-admitted.json, and
the difference between the two is load-bearing. The leave diff is taken against
admitted, not confirmed. A namespace only enters confirmed when the MDMA
confirm also succeeds, so a join that lands while its confirm is raced by the
disable itself left this node a P2P member of a namespace in neither set —
confirmed − desired was empty for it, so it was never left, and the node kept
the namespace’s key and its TEE membership indefinitely. admitted
answers the question the leave actually asks: what is this node a member of, if
it stops being entitled? — which is about P2P state, not about MDMA’s record of
it. confirmed keeps its own job, driving the confirm retry for a namespace MDMA
has not acknowledged yet.
The first boot after the admitted set was introduced
Section titled “The first boot after the admitted set was introduced”fleet-admitted.json is newer than fleet-confirmed.json, and the join loop
skips any namespace already in the confirmed set — membership is stable
until MDMA drops the assignment, so re-joining it every second would be waste.
note_admitted runs only inside that loop.
Those two facts combine badly on exactly one boot. Both files live in
/mnt/data/fleet/, so a node upgrading from a build that predates the
admitted set comes up with a populated confirmed file and an empty admitted
one — and every namespace it joined under the previous build is skipped, so
nothing ever records it. admitted − desired stays empty for those namespaces
forever, a later HA disable never makes the node leave them, and it keeps their
keys and its TEE membership: the very state the admitted set exists to
prevent, on every node that was already in the fleet.
seed_admitted_from_confirmed runs once, before the first poll, and copies
confirmed into admitted when the admitted file is absent.
Seeding from confirmed is safe in the direction that matters. A namespace
enters the confirmed set only after a local fleet-join returned
admitted=true, so confirmed is a subset of what this node is really a P2P
member of. The seed can therefore under-count — a join whose confirm never
landed appears in neither file — but never over-count, and over-counting is
the dangerous direction: it would make the node leave, and irreversibly purge
keys for, a namespace it was never admitted to.
The guard is on the file’s absence rather than its contents, because a node
that has never been assigned anything has a legitimately empty admitted set;
testing for emptiness would re-seed it on every boot and resurrect namespaces it
had since left. Once the file exists the seed never runs again, and the ordinary
prune (admitted = admitted ∩ desired) keeps it honest from there.
Two more files alongside them record what MDMA has already been
told, so nothing is re-POSTed unchanged: fleet-authorship.json for the TEE role
and fleet-inventory.json for the context inventory. Each is advanced only
after the POST it describes succeeded, so a transient MDMA blip is retried
rather than lost. Activity is logged to /run/calimero/fleet-sidecar.log.
Joining a namespace
Section titled “Joining a namespace”join_group() runs:
meroctl --home /mnt/data/calimero --node default \ --output-format json tee fleet-join <GROUP_ID>The group id is passed positionally (an earlier --group-id flag form
failed clap parsing; the positional form matches the GROUP_ID argument in the
baked core release — see the CHANGELOG for the exact core version).
Under the hood meroctl tee fleet-join generates a TDX quote, broadcasts
TeeAttestationAnnounce on the namespace topic, and polls for admission for up
to roughly 30 seconds. Critically, meroctl exits 0 whether or not the node
was admitted — process success is not admission success. So the sidecar parses
the JSON response and only treats admitted: true as success (it is tolerant of
both a flat {"admitted": ...} shape and a {"data": {...}} wrapper). If
fleet-join completes without admission, the namespace is left unconfirmed and
retried next cycle — fleet-join is idempotent in Core (an already-admitted
group returns admitted: true immediately), so retrying is cheap. This retry
convergence is what an earlier “try once, mark done” design lacked: if the group
owner was offline during the first ~30s window, the old code marked the
assignment tried and never retried.
Asking an admitter directly
Section titled “Asking an admitter directly”The broadcast only admits the node if a peer allowed to vouch — an admin, or a TEE already admitted — is in the namespace’s gossip mesh to hear it. When the only such peer is an owner’s laptop behind a relay, that mesh may form late or never, and the retry above just keeps announcing into it.
So each should-join assignment carries admitter_addrs: libp2p multiaddrs for
the fleet nodes already serving that namespace, then the owner’s node (supplied
with enable-ha). join_group() passes them as --admitter-addr flags, and core
dials each in turn and asks it over a stream, getting a verdict back instead of
a timeout. The broadcast still runs as the fallback. The addresses are dial
hints, never authority: the peer asked applies the namespace’s admission policy
and vouching rule itself.
The sidecar feeds the other half. Its poll reports swarm_addrs — merod’s own
external addresses from meroctl network status, each made to end in
/p2p/<its peer id> (a relay-circuit address names the relay’s peer first, so
“contains /p2p/” is not the test). MDMA hands those to the next node joining
the same namespace. They are cached in a file and re-read at most once a minute,
because the poll runs at 1 Hz inside a command substitution that cannot update
the parent shell’s variables.
The flags are only passed when the baked meroctl tee fleet-join --help lists
--admitter-addr. An image on a merod that predates direct admission would
otherwise fail every join on an unknown flag, which is strictly worse than the
broadcast-only join it replaces. Covered by
scripts/ci/tests/fleet-sidecar-direct-admission-test.sh.
Leaving a namespace (disable path)
Section titled “Leaving a namespace (disable path)”When a namespace the node previously confirmed is no longer in desired — HA
disabled, the slot reclaimed, or the node’s MRTD no longer trusted — the sidecar
self-leaves it:
meroctl --home /mnt/data/calimero --node default \ --output-format json namespace leave <HEX_NAMESPACE_ID>This publishes MemberLeft at the namespace root, which cascades through every
descendant subgroup where this node has a direct row. leave_group is
idempotent and non-fatal: leaving a namespace the node already left (or was
never a direct member of) returns a benign error (nothing to leave /
not a direct member) that is logged, and the loop continues — it never aborts
the sidecar.
Delegated execution: the node as a relay
Section titled “Delegated execution: the node as a relay”A fleet node does more than replicate. It is also the relay through which a Calimero Cloud account holder who runs no node at all can write.
That user holds one signing key. They cannot execute — the runtime is a WASM JIT running against materialized state under an execution lock — and they cannot decrypt, because they never received the group’s scope key. So they sign a warrant locally (their consent for one specific intent, once), present it to this node, and this node runs the method as their principal. The resulting delta is attributed to their account and device; the node’s key only signs the envelope. See Core → Delegated Authorship for the protocol.
Four facts have to reach that client before it can mint a usable warrant, and none of them is derivable anywhere but here. Reporting them is the sidecar’s second job:
| Fact | Why only the node can say it | Where it goes |
|---|---|---|
| Executor account | It is derived from an account root minted inside the TEE. MDMA never sees inside; the client has no way to compute it. A warrant naming the wrong executor is refused — after the author has spent a nonce on it. | executor_account on every should-join poll |
| Relay URL | merod sits behind Traefik on a host it was never told about, so it cannot derive its own external address — and whoever provisioned the DNS is the only party that knows the name. MDMA holds the raw public_ip, which no certificate matches. |
relay_url on every should-join poll, read from the relay-url instance metadata key |
| Relay role | Whether this node is a RelayTee (relays members’ writes) or a ReadOnlyTee (replica only) is governance state — set by the namespace’s TEE admission policy mode, which an admin of the namespace controls — readable only by a node holding that group’s key. MDMA does not set it and cannot observe it. |
authorship_ready and tee_role on /api/fleet/confirm; tee_role and per-group authorship_ready on /api/fleet/inventory |
| Context inventory | Which contexts exist inside a namespace, and which group each one hangs off, is governance state readable only by a member holding that group’s key. MDMA stores an assignment per namespace, which is one level too coarse to route a write. | namespaces[].groups[] on /api/fleet/inventory |
Only a RelayTee is authorship-ready
Section titled “Only a RelayTee is authorship-ready”Core admits an attested node as one of two TEE roles, chosen by the namespace’s
TEE admission policy (mode, set by a namespace admin):
| Role | Admitted when | Replicates, anchors sync | Relays members’ writes |
|---|---|---|---|
ReadOnlyTee |
mode: replica, or no mode (the default) |
yes | no — core answers a relay request with 403 |
RelayTee |
mode: relay |
yes | yes, by its role — no CAN_AUTHOR_ON_BEHALF bit is needed or consulted |
Both are admitted TEE members, and the sidecar treats them alike everywhere
that asks “is this node an admitted TEE here?” — the join (fleet-join’s
admitted: true is role-agnostic), the confirm, the admitted/confirmed sets,
the leave on HA disable, and the recovery-envelope walk. Only the question
“can this node execute delegated writes for users here?” separates them, and
there only RelayTee answers yes: authorship_ready is true exactly when
the node’s role in the namespace is RelayTee. A ReadOnlyTee holding a
CAN_AUTHOR_ON_BEHALF bit (from a default mask or an explicit grant) is still
not ready — core refuses it regardless, so advertising it would send clients to
mint warrants it cannot spend.
The role itself is reported alongside, as tee_role: "RelayTee",
"ReadOnlyTee", or null when this node holds no TEE role there (or a role
this sidecar does not recognise, so a role a newer core adds is never mistaken
for a relay). It is additive and optional — an MDMA that predates it ignores
it.
Why the role is re-reported, not reported once
Section titled “Why the role is re-reported, not reported once”The role is not fixed at admission. Setting a namespace’s admission policy
converts every TEE already admitted to the policy’s role, so switching an
existing namespace to relay mode — typically minutes, hours or days after this
node was admitted — turns this ReadOnlyTee into a RelayTee (and switching
back reverses it).
A join-time answer alone would therefore be frozen at false forever, and the
cloud would never advertise this relay to anyone. reconcile_authorship closes
that: each cycle it re-reads this node’s own row in the namespace root’s member
list, paged over loopback:
curl -sf "http://127.0.0.1:<SERVER_PORT>/admin-api/groups/<NAMESPACE>/members?offset=<N>&limit=1000"The list is read over loopback rather than with meroctl group members list,
because meroctl sends no offset/limit and so only ever sees core’s first
page. A partial read is never an answer: the row found on any page is certain,
but “no row” is reported only once every page has been read. Pages overlap by
one row, and that row must come back first, so a list that changed between two
pages is caught. The end is an empty page, not a short one. Each read is bounded
to MEMBER_READ_MAX_PAGES (10) pages of MEMBER_READ_PAGE_SIZE (1000, core’s
MAX_LIST_LIMIT) rows in MEMBER_READ_BUDGET (20 s). A read that fails or runs
out reports no role, which reads as not authorship-ready: the safe direction.
It re-POSTs /api/fleet/confirm only when the role changed — the loop
runs once a second and /confirm is a write — and records the new role in
fleet-authorship.json only after that POST succeeded, so a transient MDMA blip
is retried rather than lost. A conversion back to ReadOnlyTee travels back the
same way, so the cloud stops advertising the relay. A record written by an older
sidecar holds a boolean rather than a role, so it is re-reported once with the
role attached. A member list that cannot be read reports authorship_ready: false and tee_role: null — the safe direction.
For this reason MDMA’s /api/fleet/confirm accepts an already-active
assignment rather than 404’ing it: a node makes its one assigned → active
transition at admission, and requiring that state would mean a conversion to
RelayTee five minutes later could never be reported without evicting and
rejoining the node.
The context inventory, and why it is per group
Section titled “The context inventory, and why it is per group”fleet_assignments records which namespaces this node replicates. That is
not enough to answer the only question a client actually asks — can this node
execute for context X? — because a RelayTee relays in a group by its
effective role there: its row at the namespace root, which reaches a
subgroup only where membership does, and contexts belong to groups. A node can
be a relay of the namespace root while being no member at all of a restricted
subgroup beneath it.
Advertising it for that subgroup’s context makes the client discover the truth
by minting a warrant, which spends a nonce it cannot get back.
So reconcile_inventory walks the namespace’s group tree — breadth-first from
the root, since the namespace is itself a group and holds contexts of its own —
and reports each context under its own group id, with authorship_ready
re-asked for that group rather than inherited from the namespace: true only
when this node is a RelayTee in the namespace and an effective member of
the group. Membership is probed with get-capabilities, which core answers only
for an effective member; the capability bits themselves decide nothing for a
TEE:
meroctl ... group subgroups <GROUP_ID> # the walkmeroctl ... group contexts list <GROUP_ID> # what hangs off itmeroctl ... group members get-capabilities <GROUP_ID> <EXECUTOR_ACCOUNT> # membership, for a RelayTeemeroctl ... group members list <NAMESPACE> # a COUNT, never a rosterEach namespace also carries tee_role ("RelayTee", "ReadOnlyTee" or
null), read once per walk from the namespace root. It is omitted when the
role could not be read, the same rule as member_count: absent means “did not
look”, and MDMA keeps what it last recorded.
MDMA writes the result to relay_contexts, which is what
GET /api/cloud/contexts/{id}/relays reads. Without this report that table is
empty and every context answers servable: false.
The counts (member_count, context_count) ride the same report and are
aggregated on the node, inside the group’s key. No account id crosses the
wire, which is a stronger guarantee than MDMA choosing not to store one: there
is nothing there to store. A count that could not be read is omitted, never
sent as 0 — NULL means “no node has told us”, and a namespace always has at
least its admin.
On-disk bytes ride it too, as bytes: {state, private_state, delta, governance, total}, read once per scan from merod’s GET /admin-api/usage over loopback
(no token: merod runs in proxy auth mode, and traefik is not in that path).
They are RocksDB estimates, so they are excluded from change detection —
counting them would post on every scan — and refresh on whatever post happens
next, at worst the 15-minute full pass. An unreachable /usage omits bytes
and changes nothing else about the report; bytes never block a prune.
On change, with a slow full reconcile
Section titled “On change, with a slow full reconcile”Core exposes no “contexts changed” signal, so the walk is polled: once a minute,
posting only when the payload differs from the last post MDMA acknowledged. A
steady-state node therefore posts nothing at all. Every 15 minutes, and
immediately whenever the confirmed set changes, it posts with full: true.
Both halves are load-bearing. Walking at the loop’s 1 Hz would be dozens of
meroctl subprocesses a second for data that moves a few times a day; and an
incremental post can only add, so without the periodic full pass a context the
node has left would be advertised forever.
A partial read is never posted as authoritative
Section titled “A partial read is never posted as authoritative”full: true makes MDMA prune every context row the report does not mention,
per namespace. So a full report assembled from partial reads deletes rows for
contexts this node is serving perfectly well, and un-advertises a healthy relay
because one meroctl call happened to fail.
The walk therefore tracks whether it saw everything, and anything that could
hide a context clears that flag: a failed or unparseable read, or a list that
came back at core’s 100-row page limit — meroctl passes no offset/limit, so a
full page is indistinguishable from a truncated one. Namespaces that came back
complete are posted with the requested full; the rest are posted additively in
a second request, and MDMA keeps what it already has for them.
The distinction is deliberately narrow: only a read that could hide a context
counts. A failed role, membership or member-count read is a worse answer about
contexts that are still being reported, so it degrades to false/absent without
blocking the prune — otherwise a namespace whose member list merely exceeds one
page could never be pruned at all.
Attested registration, and why nothing here signs anything
Section titled “Attested registration, and why nothing here signs anything”Everything the sidecar tells MDMA about itself — peer id, executor account, relay URL, MRTD — was taken on trust. The only credential is a fleet-wide shared token, which identifies some node and never which node, so MDMA could not tell one node’s claims from another’s. Two things rested on that:
should_joincompared the reported MRTD against a customer’sallowed_measurements. That value is a hex string the caller types, so the allowlist meant to restrict a namespace to nodes running a known image was satisfied by asserting the right answer.- relay identity was keyed on the claimed peer id, and
relay_url/executor_accountare whatGET /api/cloud/contexts/{id}/relayshands clients — so naming yourself the relay for a context bought denial of service and disclosure of intent contents.
GET /api/fleet/nodes/challenge → POST /api/fleet/nodes/register closes both.
See mdma#225.
No signature, and no core change was needed. The executor account is minted
inside the TEE and never leaves it, so an allowlisted image saying what its
account is is precisely what attestation is for. The binding rides the quote’s
report data: the sidecar computes
SHA256(challenge | peer_id | executor_account | relay_url), asks merod’s
merod for a quote over that nonce (see below), and MDMA recomputes
the same hash from the request body. A quote therefore commits to exactly one
identity and one challenge — a genuine quote from one node cannot be presented
alongside another’s claim, and neither can be replayed.
The quote comes from POST /admin-api/tee/registration-attest, on core’s
protected router, which the sidecar reaches over loopback (these nodes run
merod in proxy auth mode, so traefik is the only guard, and it does not expose
the route). That route fills the second half of report data with
attest_registration_binding(), a value the public /admin-api/tee/attest
never produces, and the sidecar sends quote_binding: "registration" so MDMA
checks the quote against it. A merod that predates the route answers 404; only
then does the sidecar fall back to /admin-api/tee/attest with nothing bound
(second half zeros) and send no quote_binding. MDMA accepts that shape until
fleet_registration_require_binding is turned on, which is safe once every
fleet node runs an image with the route. traefik-tee-route-test.sh pins the
public TEE router to exactly info and attest.
Four things this has to get right, each covered by
scripts/ci/tests/fleet-sidecar-registration-test.sh:
- The hash is a contract computed twice. Field order and separator are part of it; get either wrong and every registration is refused as bound to a different identity — indistinguishable from a genuinely bad attestation.
- Only an accepted registration is recorded. State is written after MDMA says yes, so a refusal or an outage is retried rather than remembered as done.
503is not403. MDMA answers503when the quote could not be evaluated — Intel collateral unreachable, verifier missing — which is not a verdict about this node. Treating it as a refusal would leave a healthy fleet unregistered through someone else’s outage.- Re-register on change, not on a timer. A registration does not expire, and the identity it binds is fixed for the life of the node unless it re-keys or moves behind a different ingress.
Until a node registers it replicates but does not relay: MDMA records no identity for it and never advertises it, which is the same posture every other failure on this path takes.
Account recovery envelopes, and who they are written for
Section titled “Account recovery envelopes, and who they are written for”Holding an account root recovers the account. It does not recover the list of namespaces that account belongs to: namespaces are separate keypairs, and the local governance state left with the lost device. Without a record somewhere, those namespaces keep existing on other people’s nodes and are permanently unaddressable — and the loss is only discovered at the moment the last device dies, when no retroactive fix exists.
A relay is the only party that can write that record. It holds the group key, so
it can enumerate the membership, and since
core#3918 it can seal to a
member’s account root via meroctl account seal-to. The recipient being the
root rather than a device key is the whole point: a device is exactly what is
gone in the case worth sealing for, so an envelope addressed to one is valid,
opaque, and unopenable precisely when it is needed.
One envelope per (account, namespace) this node serves, never one per
account. A relay only ever sees the namespaces it serves, so a per-account
envelope would be a partial list overwriting other relays’ partial lists — which
is mdma#228, fixed by
keying MDMA’s table per namespace.
What the sealed body carries
Section titled “What the sealed body carries”{ "v": 2, "namespace_id": "<64 hex>", "groups": ["<64 hex>", "..."], "contexts": ["<64 hex>", "..."], "devices": [{ "device_id": "<hex>", "signing_key": "<64 hex>" }]}namespace_id is repeated inside the sealed body so an envelope is
interpretable without trusting the cleartext key it arrived under.
devices holds only the recipient’s own devices, deduplicated across the
groups they appear in and sorted. It is there because device_id is what
POST /namespaces/:id/account/revoke takes, and a holder who lost every device
cannot obtain it from a node afterwards:
GET /admin-api/account/devices is unmapped in core’s permission validator, so
it falls to default-deny and needs admin, while the session a recovering
holder can actually get is account_proof — context:query, context:intent
and context:subscribe, and deliberately nothing else. The envelope is the only
path, so the list is sealed into it rather than looked up later.
Another member’s device ids are not carried. They are no use to a recovering holder, and would widen what a single envelope discloses about everyone else in the namespace — the opposite of what sealing per account is for.
Four things this writer has to get right, none of which fail loudly:
- Every member, linked or not. This used to write only for accounts holding
a cloud login, on the grounds that nobody else could read an envelope: the
only read path went through
/api/cloud/me/accounts/{id}/recovery-envelope, which needs a session and a link. That premise is gone — MDMA now servesPOST /api/cloud/accounts/recovery-envelopeagainst a root signature, so a keyholder can read its own envelopes having never seen a cloud login. Filtering would mean an unlinked member who loses their last device has nothing to recover from, and finds that out at the one moment no retroactive fix exists. The cost is storage, and MDMA learning the(account, namespace)edge for non-customers — an edge it could already derive, since it assigned the relay that would write the row. - Incomplete means write nothing. A shorter list overwriting a longer one at a higher version is the same data loss as a stale clobber, with a valid version number on it. Any failed read in the walk aborts that namespace’s pass.
- The version is durable. It lives in
/mnt/data/fleet/fleet-recovery.json, not in memory and not in wall-clock. MDMA answers409to a version that goes backwards, so a counter that reset on restart would strand the envelope at whatever MDMA last accepted. - Change is detected on the plaintext. Sealing mints a fresh ephemeral sender key every call, so identical data seals to different bytes. Comparing ciphertexts would see a change on every pass, burn a version each time, and rewrite MDMA forever for data that never moved.
Members who leave
Section titled “Members who leave”An envelope is the only (account, namespace) edge MDMA holds, and MDMA lists a
namespace in the cloud console of every login whose account has one. So an
envelope that outlives the membership keeps showing the namespace to someone who
was removed from it.
After a complete read, the sidecar removes the envelope of every account it
wrote for in that namespace that the read no longer lists
(DELETE {MDMA_URL}/api/fleet/recovery-envelope), then drops that account from
its state file. Two limits keep this from deleting anything it should not:
- Only after a complete read. The same bar as writing: a read that failed part-way is not evidence that anyone left.
- Only accounts this node wrote for before. A relay that has not caught up on a join has no entry for the new member, so its out-of-date view cannot remove an envelope a better-synced relay wrote.
Dropping the local entry means a member who later rejoins starts again at
version 1. MDMA bounds a new row’s version, so carrying the old counter forward
could earn a 409 on that first write.
MDMA’s ciphertext field is a single opaque string, while a sealed envelope is
four values (accountRootEpoch, ephemeralPublicKey, nonce, ciphertext) and
a recovering holder needs all four. The whole envelope is JSON-encoded into that
string; MDMA stores it without looking, as it does everything else here.
Every failure lands on “replicates but does not relay”
Section titled “Every failure lands on “replicates but does not relay””None of these reads is fatal, and they degrade in the same safe direction.
A node that cannot read its account, has no relay-url metadata, fails the
role or membership read, or cannot walk its group tree reports the missing value as
empty/false/absent. MDMA then does not
offer it as a writable relay, so clients are never sent to mint warrants it
would refuse — they see “no relay available” rather than a refusal after
burning a nonce. The node keeps replicating throughout.
An empty report never erases a value MDMA already holds (only a different non-empty one overwrites), which is what lets an older and a newer sidecar be live at once during a fleet roll.
The edges the image opens
Section titled “The edges the image opens”Being a relay means accepting a write from a caller with no account on this
node, so the image has to open exactly three paths — and each carries its own
credential. merod refuses an intent, before executing anything, unless
the warrant is genuinely the member’s, commits to this exact context, method and
arguments, has not expired, has an unspent nonce, and this node is a RelayTee
in the owning group. A node token would prove none of that, and
requiring one would make the feature unreachable for exactly the callers it
exists for.
All three edges are driven by the single Ansible variable fleet_delegated_access
(defined at play level in mero-tee/playbook.yml; on for the ReadOnly fleet
profiles, off for debug):
| Edge | What it does |
|---|---|
Traefik (traefik-routing.yml.j2) |
Adds a node-api-intents router for PathRegexp(^/admin-api/contexts/[0-9a-f]{64}/intents$), GET/POST/OPTIONS only, at a higher priority than the auth-guarded node-api catch-all and without the auth-node forwardAuth middleware. A PathPrefix on /admin-api/contexts/ would have exempted every context route on the node, DELETE included. |
Traefik (node-api-context-intents) |
Adds the same kind of router for delegated context creation: PathRegexp(^/admin-api/groups/[0-9a-f]{64}/context-intents$), GET/POST/OPTIONS only, same priority, rate-limit tier and CORS, no auth-node. The body holds a creation warrant signed by the author’s device key; merod refuses it unless the author holds CAN_CREATE_CONTEXT (or is an admin) in the group and this node is a RelayTee there (or holds CAN_AUTHOR_ON_BEHALF). GET ...?author=<hex> is the descriptor; the query string is not part of what Traefik matches. A PathPrefix on /admin-api/groups/ would have exempted /admin-api/groups/<group>/contexts and every membership route. |
Traefik (node-api-governance-intents) |
Adds the same kind of router for delegated governance: PathRegexp(^/admin-api/groups/[0-9a-f]{64}/governance-intents$), GET/POST/OPTIONS only, same priority, rate-limit tier and CORS, no auth-node. The body holds a governance warrant signed by a group member’s device key, which merod verifies before applying the change. Every other group route, /admin-api/groups/<group>/members included, stays behind auth-node. |
merod (calimero-init.sh.j2) |
Sets server.admin.delegated_access=true, which moves the intents, context-intents and governance-intents routes onto core’s public router. |
Traefik (node-sealed) |
Routes POST /sealed/v2/handshake and POST /sealed/v2 (exact paths, plus OPTIONS) to merod, also without auth-node: forwardAuth cannot see the route a sealed request names, which is inside the envelope. See sealing the intent below. |
The same variable also turns on device-key login, scoped per tenant: the
sidecar reads this node’s device signing key from GET /admin-api/identity,
records it for mero-auth-start and restarts mero-auth when it changes, and
merod takes each session’s account from the proxy (server.proxy_identity).
See device-key login.
They are one variable rather than two defaults on purpose: these nodes run
merod in proxy auth mode, so Traefik is the only gate. An exemption added
there without the matching merod setting still opens the path — the drift is
silent in the unsafe direction.
Sealing the intent to the TD
Section titled “Sealing the intent to the TD”TLS from a browser ends at Traefik, on the host, outside the TD. Everything in an
intent (the warrant and the method’s arguments) is plaintext there. With the
node-sealed router, a client can instead attest merod’s transport key
(POST /admin-api/tee/attest, already public) and seal the intent to it: Traefik
forwards opaque POST /sealed/v2 requests, and only merod, inside the TD,
opens them. mero-js does this with connectCloud({ seal }).
The router is exempt from auth-node because forwardAuth could not apply to it,
and that is safe only because merod enforces the boundary itself: in proxy auth
mode it lets an opened request reach nothing but what it serves without a
credential, which here is the intents, context-intents and governance-intents
routes and the public probes and TEE routes. Which inner paths qualify is decided by merod, not by
this repository: there is no inner-path allowlist here to keep in step.
A sealed request for anything auth-node guards is refused inside the envelope
with 403 sealed_route_unguarded. So on these nodes intents are sealed, while
logins (/auth/, served by mero-auth, outside merod) and session-bearing reads
are not.
merod enforces that rule from 0.11.0-rc.52. An older merod routes an opened
request anywhere, which would make the router a way around every guard on the
node, so scripts/ci/tests/traefik-sealed-route-test.sh refuses the router with
an older merodVersion. It also pins the router to the relay-only block, to the
two exact paths, and to the tight rate-limit tier.
TLS: the certificate a browser needs
Section titled “TLS: the certificate a browser needs”A relay URL is useless to the audience delegated execution exists for unless it
is https://. The client is a page served over HTTPS, and a page served over
HTTPS may not call an http:// origin — mixed content is blocked outright,
with no user override and nothing the site can opt into. https://<public-ip>
does not help either: no publicly-trusted certificate matches a raw address that
the fleet re-assigns on every stop/start.
So the node terminates TLS itself, on Traefik’s websecure entrypoint, and the
sidecar is what gets a certificate onto it.
The key never leaves the enclave
Section titled “The key never leaves the enclave”calimero-init generates an EC P-256 key and a self-signed placeholder at
/mnt/data/tls/ before Traefik starts, so the TLS store always loads. What
travels to MDMA is a CSR over that key; what comes back is a signed
certificate. MDMA holds no private key for any node, and a compromise of its
certificate table leaks only material already public in Certificate
Transparency.
The placeholder deliberately reports no expiry to MDMA. It is valid for ten years, and reporting that would tell the control plane this node is covered until 2036 — no certificate would ever be ordered, and the node would serve an untrusted certificate forever with nobody told.
Why Traefik has no ACME resolver
Section titled “Why Traefik has no ACME resolver”This is the constraint that shapes the whole design, and it is invisible from outside the role.
Traefik’s own certificate resolver needs this node’s hostname in its config —
a certificatesResolvers block, tls.domains, or a Host() rule on each
router. traefik-routing.yml is rendered into the image, so a per-node name
there changes the image’s root hash and therefore its measurements. A fleet
whose MRTD differs per machine cannot be allowlisted, which is the entire point
of measuring it.
The rule that survives is narrower than “no certificates on nodes”:
:80 stays open beside :443 for the same reason it always was: a CLI, the
desktop app and merobox reach a node by raw IP and need no certificate.
Issuance is authorised by two independent facts
Section titled “Issuance is authorised by two independent facts”MDMA orders the certificate over ACME DNS-01, which means proving control of the name by writing a DNS record — the node is never contacted, so a certificate can be issued for a node that is firewalled, still booting or briefly down. Two things have to hold before it will order one:
- The CSR’s digest is inside the attestation quote. It joins the
registration nonce as a fifth field —
SHA256(challenge || peer_id || executor_account || relay_url || sha256(csr))— so the quote commits to the key being certified, not only to the identity asking. Without it, a quote is a valid identity claim that says nothing about the key travelling beside it. - The hostname is recomputed by MDMA, from the node’s own record and the
deployment’s relay template — never taken from the
relay_urlthis node reports. A node chooses what it reports and its quote is genuine either way, so the binding alone cannot stop it asking for someone else’s name.
A CSR failing either check leaves the node registered and replicating. It asked for something it may not have; it is not an impostor.
The four-field binding is still accepted, because that is what a node image from before node TLS sends and what a node with no relay URL sends today. Refusing it would unregister the fleet to roll out a feature that fleet does not use.
Delivery and renewal ride the poll
Section titled “Delivery and renewal ride the poll”No new endpoint and no blocking call. Issuance is asynchronous on MDMA’s side —
it has to order and wait for validation — so the chain simply appears in the
tls_certificate field of whichever should-join response follows it. Renewal
takes the identical path: the poll reports tls_cert_not_after, MDMA orders one
for a node reporting none and renews one whose expiry is near, and the chain is
returned only while the two disagree.
Before installing anything, the sidecar checks the returned chain’s public key against the key it holds. A mismatched certificate fails nowhere visible — Traefik loads it happily and every client handshake breaks, on a node MDMA has already advertised as a relay.
Traefik is then restarted, not reloaded: it watches its dynamic config file, not the certificate files that file names, and the config sits on the measured read-only root where it must not be rewritten to provoke a reload. Identical bytes are a no-op, so a poll that runs every second does not restart it every second.
Metrics carry the node’s identity
Section titled “Metrics carry the node’s identity”vmagent scrapes localhost:9100 (node_exporter listens on 127.0.0.1:9100, and vmagent’s own HTTP listener on 127.0.0.1:8429: both are loopback only), which is the same target string on every node
in the fleet. Without identifying labels every node therefore produces the
identical series identity:
{instance="localhost:9100", job="node_exporter"}Two nodes and their samples interleave into one series — a counter that appears to reset, a gauge that flips between machines. That reads as plausible data and is wrong, which is worse than having no metrics at all. With a single node in the fleet it is invisible; the second node starts corrupting every query silently.
configure_vmagent.sh therefore starts vmagent with:
-remoteWrite.label=instance_name=<fqdn>-remoteWrite.label=instance_type=meroteeDerived inside that script rather than passed in, so its two callers cannot
drift: it runs from calimero-init at boot and from the fleet sidecar when
MDMA rotates the observability token, and a rotation that dropped the labels
would orphan the node’s metrics from that moment on.
instance_name and instance_type are the same labels the Terraform-managed
nodes already emit, so one dashboard covers both kinds. instance_name is the
FQDN, which is exactly the value Vector writes into the logs’ hostname field —
that is the join between a node’s logs and its metrics.
Both signals also carry which image the node runs, as instance_profile:
| metrics | -remoteWrite.label=instance_profile=<profile> |
| logs | .instance_profile, and listed in VL-Stream-Fields so it is indexed |
locked-read-only is production. The debug profiles are rebuilt freely and
their certificates come from Let’s Encrypt staging, so they are not publicly
trusted — and directory_for_profile() keys that off the profile rather than a
deployment setting, so a debug node physically cannot spend the production CA’s
issuance budget. Without the field, telemetry from a throwaway node reads
exactly like production’s and the only way to tell them apart is to
cross-reference nodes.image_profile in MDMA’s database, which nobody does
before believing a graph.
Both read /etc/calimero/image-profile, which the playbook writes at build time
and the measured root hash covers, so a node cannot misreport it. Vector’s VRL
cannot read a file, so the partial config carries a __IMAGE_PROFILE__
placeholder that configure_vector.sh substitutes — a placeholder that shipped
unsubstituted would make every node report that literal string, which is what
its test pins. A missing file yields unknown, never an empty value: blank
would silently drop the dimension for that node alone while every other node
still reported one.
Covered by scripts/ci/tests/vmagent-node-identity-test.sh and
scripts/ci/tests/vector-image-profile-test.sh, both of which run the real
scripts and read what they generate.
Registration runs before the poll, and must
Section titled “Registration runs before the poll, and must”/api/fleet/should-join requires a fleet token. A node is issued one by
registering, and /api/fleet/nodes/register is a bootstrap route that takes
none. So the order of those two inside the reconcile loop decides whether a
fresh node can join the fleet at all:
poll first 401 -> `continue` -> registration never runs -> no token -> poll 401s again, forever. A closed loop.
register first token issued -> poll succeeds -> everything else runs.The closed version is also silent. poll_mdma returns 1 with no output by
design, so a node stuck in it logs nothing after its startup lines while looking
healthy from every outside angle: merod serving, correct image, MRTD in the
published allowlist, metadata stamped. It replicates, and it is never advertised
as a relay, and nothing says why.
This was unreachable while the image carried a baked fleet token. FLEET_TOKEN
was already set at boot, the first poll succeeded, and the loop fell through to
registration. Removing that secret from the image — which was right, and is what
fleet-sidecar-logs-token-test.sh exists to protect — left the ordering behind,
and every node created afterwards stopped registering.
Pinned by scripts/ci/tests/fleet-sidecar-bootstrap-order-test.sh, which
asserts the loop line of reconcile_registration precedes the should-join
gate, that reconcile_executor_account precedes it in turn, and that the gate
still continues — the safety property below has to survive the reorder.
The poll failure is no longer silent either: it logs once on a state change, distinguishing “no token yet” (bootstrapping, or registration is being refused) from “token rejected” (mdma rotated its signing key, and re-registration will issue a new one).
The poll-success safety gate
Section titled “The poll-success safety gate”The step-2 gate is safety-critical, not a nicety. Because a missing assignment
now triggers an irreversible namespace leave + key purge, a transient MDMA
outage must never be mistaken for “all namespaces disabled” — that would shred
keys on every healthy replica at once.
So poll_mdma captures curl’s body and exit status separately and returns:
0+ a validated body only on HTTP 200 whose body parses as an{"assignments": [...]}object in which everygroup_idis a 32-byte hex namespace id; or1with no stdout on any curl error, non-2xx, timeout, or unparseable/wrong-shape body, including a single malformedgroup_id.
A malformed id fails the whole poll rather than being dropped: dropping it would make that namespace look unassigned and trigger its leave.
On a 1, the main loop skips the entire reconcile — no desired computed, no
join, no leave, no prune — and preserves fleet-confirmed.json and
fleet-admitted.json untouched.
Preserving confirmed across a failed poll is itself load-bearing: if a failed
poll pruned the sets, a genuine disable on the next good poll would no longer
appear in the admitted − desired diff, and the node would never leave.
Symmetrically, confirm_assignment returns curl’s real exit code (rather than
swallowing errors), so a transient blip on the confirm step leaves the namespace
out of confirmed and is simply retried next cycle instead of being lost.
MDMA endpoints and paths at a glance
Section titled “MDMA endpoints and paths at a glance”| Item | Value |
|---|---|
| Poll (should-join) | POST {MDMA_URL}/api/fleet/should-join — body {"peer_id", "mrtd", "executor_account", "relay_url", "tls_cert_not_after"}; response may carry tls_certificate |
| Confirm | POST {MDMA_URL}/api/fleet/confirm — body {"peer_id", "group_id", "authorship_ready", "tee_role"}; tee_role is "RelayTee", "ReadOnlyTee" or null, and authorship_ready is true only for RelayTee |
| Join command | meroctl … tee fleet-join <GROUP_ID> [--admitter-addr <MULTIADDR>]… (positional group id) |
| Leave command | meroctl … namespace leave <HEX_NAMESPACE_ID> |
| TEE role probe | GET http://127.0.0.1:{server-port}/admin-api/groups/<NS>/members?offset=&limit=1000 (loopback, paged) → this node’s members[].role |
Group membership probe (inventory, RelayTee only) |
meroctl … group members get-capabilities <GROUP_ID> <EXECUTOR_ACCOUNT> succeeds |
| Readiness probe | meroctl --output-format json peers |
| PeerId source | /mnt/data/calimero/default/config.toml → [identity].peer_id |
| MRTD source | /sys/class/misc/tdx_guest/measurements/mrtd:sha384 |
| Executor account source | meroctl --output-format json account show → accountId |
| Relay URL source | Instance metadata relay-url (written by MDMA’s dispatcher from MDMA_TEE_RELAY_URL_TEMPLATE) |
| Confirmed state | /mnt/data/fleet/fleet-confirmed.json |
| Reported TEE role state | /mnt/data/fleet/fleet-authorship.json — {<namespace>: "RelayTee" | "ReadOnlyTee" | null} |
| Node challenge | GET {MDMA_URL}/api/fleet/nodes/challenge → {"challenge", "expires_at_ms"} |
| Node registration | POST {MDMA_URL}/api/fleet/nodes/register — body {"peer_id", "executor_account", "relay_url", "challenge", "quote", "csr"}; returns {"fleet_token"} |
| Quote source | POST http://127.0.0.1:{server-port}/admin-api/tee/registration-attest (on 404, /admin-api/tee/attest) — body {"nonce"}, loopback; server-port must be 1–65535, else 2428 |
| Registration state | /mnt/data/fleet/fleet-registration.json |
| Recovery membership walk | meroctl … group members list | group subgroups | group contexts list | group member-devices <GROUP_ID> |
| Recovery envelope | PUT {MDMA_URL}/api/fleet/recovery-envelope — body {"peer_id", "account_id", "namespace_id", "ciphertext", "version"} |
| Recovery envelope removal | DELETE {MDMA_URL}/api/fleet/recovery-envelope — body {"peer_id", "account_id", "namespace_id"} |
| Seal command | meroctl … account seal-to <GROUP_ID> <ACCOUNT> (plaintext on stdin) |
| Recovery envelope state | /mnt/data/fleet/fleet-recovery.json |
| TLS key and certificate | /mnt/data/tls/key.pem, /mnt/data/tls/cert.pem — key generated by calimero-init before Traefik starts; never leaves the node |
| TLS reload | systemctl restart traefik (it watches its config file, not the certificate files that file names) |
| Log | /run/calimero/fleet-sidecar.log |
| Poll interval | 1 second |
| Auth header | X-Fleet-Token, once registration has issued one; manager signs with MDMA_FLEET_TOKEN_SIGNING_KEY. The two nodes/* bootstrap routes take none. |
Where to go next
Section titled “Where to go next”- System overview — the two trust flows, and where the sidecar sits in fleet admission.
- Key release flow — the other key mechanism (KMS storage key), kept deliberately separate from fleet admission.
- Trust model — what MRTD and the other measurements pin.
- Glossary — MRTD, MDMA, namespace,
ReadOnlyTee, and the rest of the vocabulary. - Operate: runbooks and release pipeline — how the node image carrying this sidecar is built and rolled out.