Adding a storage medium to UMBP distributed mode
A backend owns bytes and publishes descriptors for them. It does not move
them. Everything in this skill follows from that one rule, which is stated at
the top of src/umbp/include/umbp/distributed/peer/backend/medium_backend.h
and enforced by the type system since Phase 6 (Init receives a
MemoryRegistrar, which has no Submit on it).
If you find yourself wanting to copy bytes inside a backend, stop — that is the
transfer layer's job, and the sibling skill umbp-add-transfer-engine covers
it.
First decide which of the two shapes you need
This is the whole design decision, and getting it wrong costs a rewrite.
Is your medium page-addressable — can you hand out a raw pointer that the transfer layer can read and write directly?
| Yes (DRAM, HBM, CXL, a pmem mapping) | No (SSD, S3, a network filesystem) | |
|---|---|---|
| Implement | PageMemorySource — 5 methods |
MediumBackend — 23 methods |
| Reuse | All of PageBackend |
Nothing; but see "staging" below |
| Effort | ~100 lines | ~450 lines |
| Example | hbm_backend.{h,cpp} |
ssd_backend.{h,cpp} |
Most media are the first case. PageBackend's slot lifecycle, bitmap
allocator, event outbox, reaper, read leases and copy pins never consult the
tier, so a second paged medium needs no second copy of any of it.
Shape 1: a paged medium (implement PageMemorySource)
Five methods, in page_backend.h:
class PageMemorySource {
virtual bool Allocate(const std::vector<uint64_t>& sizes, std::vector<Buffer>* out) = 0;
virtual void Release() = 0;
virtual mori::io::MemoryLocationType LocationType() const = 0;
virtual int Device() const = 0;
virtual const char* Name() const = 0;
};
LocationType() and Device() are the facts a descriptor cannot recover.
They are the reason this interface exists: PageBackend mirrors them into every
TransferRef it publishes, and that is what makes a local transfer against your
medium select the right engine. Get them wrong and nothing fails at build or
init time — a transfer just silently picks the wrong engine, or no engine at
all.
The recipe
Write the source. Put it in its own file if it needs a new dependency (
hbm_backend.cppexists so HIP stays out ofpage_backend.cpp).bool MyPageMemorySource::Allocate(const std::vector<uint64_t>& sizes, std::vector<Buffer>* out) { std::vector<MyHandle> taken; std::vector<Buffer> staged; for (uint64_t size : sizes) { if (size == 0) continue; // skip, do not fail MyHandle h = MyAlloc(size); if (!h.valid()) { for (auto& t : taken) MyFree(t); // unwind THIS call only return false; // leave `out` untouched } staged.push_back(Buffer{h.ptr, h.usable_size}); taken.push_back(h); } handles_.insert(handles_.end(), taken.begin(), taken.end()); out->insert(out->end(), staged.begin(), staged.end()); return true; }Report the usable size in
Buffer::size, not the requested one — host hugepage rounding makes the extra genuinely allocatable, andPageBackendpublishes what you report.Add a factory returning
std::unique_ptr<MediumBackend>, next toMakePageBackend/MakeHbmBackend. It must return the interface: Phase 5 Rule A says onlyPoolClient::Initmay name a concrete backend.std::unique_ptr<MediumBackend> MakeMyBackend(uint64_t page_size, ...) { return std::make_unique<PageBackend>(TierType::MY_TIER, page_size, std::make_unique<MyPageMemorySource>(...), std::move(buffer_sizes), pending_ttl, read_lease_ttl); }Add a config struct in
distributed/config.hholding your medium's own knobs — noenabledflag:PoolClientConfig::mediumis the single selector. Do not extendDramOwnershipConfig— hugepages/NUMA/prefault are meaningless for HBM, and a device ordinal is meaningless for host memory. Each medium brings its own knobs; that asymmetry is why the seam is a class and not an options struct.Add a case to the medium switch in
PoolClient::Init. A node registers exactly one backend, chosen byconfig_.medium, so you add a case rather than anotherif. This is the only file that changes outside your own:case TierType::MY_TIER: { backend = MakeMyBackend(page_size, config_.my, ...); break; }The shared
Init/Register/ error handling below the switch is written once for every medium.Why one medium, not a tier stack. The routing plane does not tier: master treats every advertised tier as an equally valid put target (Phase 4 deleted the hardcoded tier orders), so a node registering two backends mirrors across them rather than promoting/demoting between them. Heterogeneity comes from different nodes picking different media. If you find yourself wanting two live backends on one node, you are asking for a local tiering policy that does not exist yet — that is a routing-plane change, not a backend one.
Add the tier to
TierType(types.h) if it is genuinely new.HBM=1, DRAM=2, SSD=3already exist.BackendRegistryis amap<TierType, ...>, so one instance == one medium and the enum value is the identity. Also add a matchingUMBPMediumvalue incommon/config.hand map it inToTierType, or no user-facing config can select your medium.
That is the whole change. Routing, the peer service, the heartbeat and the batch
executors were all written against BackendRegistry and need no edit.
Metrics: do not write any
PoolClient::Init wraps whatever the medium switch produced in
InstrumentedBackend before registering it, and that decorator derives the
whole generic series — operations by outcome, bytes committed / resolved /
freed, batch depth, time spent inside the medium — from the MediumBackend
calls themselves. Your backend is charted the moment it is composed in, under
tier=<your tier>, in the panels of
examples/monitoring/grafana/dashboards/umbp_backends.json that already exist.
There is no metric to register, no dashboard to edit, and no counter to add to
your backend.
Override SampleMetrics() (from MetricSource) only for state the interface
cannot show from outside — what the device itself did, how full an internal
arena is. When you do, obey the one rule that keeps the single dashboard
working: publish under the generic name and put your specifics in a label.
std::vector<MetricSample> MyBackend::SampleMetrics() const {
return {MetricSample{MORI_UMBP_METRIC_BACKEND_MEDIUM_EVENTS_TOTAL,
MORI_UMBP_METRIC_BACKEND_MEDIUM_EVENTS_TOTAL_HELP,
{{"event", "device_read_error"}}, // YOU name the event
device_read_errors_.load(std::memory_order_relaxed)},
MetricSample{MORI_UMBP_METRIC_BACKEND_MEDIUM_STATE,
MORI_UMBP_METRIC_BACKEND_MEDIUM_STATE_HELP,
{{"state", "queue_depth"}},
depth, MetricKind::kGauge}};
}
A metric named after your medium (mori_umbp_myssd_reads_total) compiles and
scrapes fine, and is still wrong: it needs its own panel, which is how UMBP
ended up with a dashboard per medium and an SSD dashboard wired to counters
that lost their publisher. tier= and backend= are stamped by the publisher —
do not set them. Counter values must be monotonic; gauges may move either way.
See umbp/distributed/metrics/component_metrics.h.
Shape 2: a medium whose bytes are not addressable
If you cannot hand out a pointer, you have two options, and the cheap one is almost always right.
Stage through registered host memory (what SsdBackend does)
Publish an ordinary registered host DRAM arena as your buffer, and move bytes between it and your device inside the backend:
BatchAllocatereserves a staging page and publishes it. The writer RDMAs into it knowing nothing about your medium.BatchCommitspills that page to your device and returns the page to the arena. The page is borrowed for the write, not the key's home.BatchResolvefills a staging page from your device and publishes it under a read lease; a reaper reclaims the page when the lease expires.
The cost is one host copy per side. The benefit is that your medium reaches the
data plane with zero new transfer-layer concepts — no new TransferRef
kind, no new engine, no chaining — and a remote peer reading from your node sees
a perfectly ordinary registered buffer and needs no code at all.
Add a real endpoint kind (bigger; not yet done in tree)
transfer_engine.h reserves this: a FileRef/ObjectRef plus a kind tag, a
new engine, and chaining in CompositeTransferEngine for the remote reader
(device → bounce → wire), which that class explicitly does not implement. Take
this path only when you need zero-copy/GDS; staging does not block it.
Things that bite in shape 2
- Exhaustion has no honest encoding. A
Resolvethat cannot get a staging page must reportfound=false, which makes the client exclude your node and retry elsewhere — wrong, because your node does hold the key.medium_backend.hrecords that the "not ready, retry here" state was proposed and rejected. Size the arena for read concurrency, and pin the behavior in a test so a future control-plane fix changes it deliberately. - Contiguity.
PeerSsdManager::PrepareReadtakes one(ptr, capacity), not a scatter list, so a staged key must fit one contiguous page.SsdBackendtherefore allocates one buffer ofstaging_pages * page_sizeand refuses keys larger than a page. In distributed mode master'spage_sizeis the KV block size, so 1 key == 1 page is the normal case. - Never hold the arena lock across IO. Take the slot out under the lock,
release it, do the spill/fill, then re-lock to return the page. Holding it
stalls every concurrent
AllocateandResolve.
The contracts that are easy to get wrong
These are the ones with no compile-time protection.
Events go in ONE bundle under ONE seq. The heartbeat concatenates every
backend's events (DrainAllBackends). Never emit one bundle or one seq per
medium — that breaks the ack / seq-gap full-sync recovery.
SnapshotOwnedKeysForFullSync must clear the outbox in the SAME critical
section as the snapshot. Two separate locks drop events committed in between.
SnapshotOwnedKeys is const and must not mutate; the full-sync variant is
not.
kFailedNoSpace vs kFailed. kFailedNoSpace means "medium exhausted,
retry on another peer". Use it only when another peer could plausibly succeed. A
key too large for your page size is kFailed — no peer would do better, and the
writer must not keep hunting.
Evict returns one result per key, in request order. The peer service sums
freed bytes for a key mirrored across media and relies on the positional match.
Return bytes_freed = 0 (not an error) for unknown, already-freed, or protected
keys; master retries protected ones next round.
BatchAbort is idempotent — a slot already reaped or never seen still
reports true.
Shutdown must tolerate a failed Init and must not run concurrently with
anything else. Deregister before releasing memory: the registrar may still
hold an MR over those pages.
Init is idempotent, and a backend is fully live when it returns — start
your own reaper thread there, so no caller has to know you have one.
Allocation is gated between ClearLocal() and ClearFullSyncAcked(). No
new owned key may appear in that window.
Threading: every method may be called from the peer service's gRPC handler threads and the heartbeat thread concurrently.
SetAutoFlushHook's callback runs under your lock — it must be cheap and
must not re-enter the backend. Signal the heartbeat thread and return.
Testing
Two levels, both worth having:
Without the transfer layer — drive
BatchAllocate/BatchCommit/BatchResolve/Evictdirectly and assert the bookkeeping. ALocalOnlyRegistrar(returnsTransferRef::HostBytes, counts register/deregister calls) is all theMemoryRegistraryou need; it is also exactly whatCompositeTransferEnginedegrades to on a node with no RDMA.Through a real
CompositeTransferEngine— a Put and a Get moving actual bytes. This is what catches a wrongLocationType()/Device(), because a mislabeled endpoint selects the wrong engine and the copy either fails or silently does nothing.
Assert the registration contract explicitly:
EXPECT_EQ(registrar.last_loc, mori::io::MemoryLocationType::GPU);
EXPECT_EQ(backend->BufferRef(0).device, 0);
For a staging backend, the highest-value test is that Commit returns the
page: configure 2 staging pages, do 10 sequential puts, assert all succeed. A
leaked page wedges the backend after staging_pages puts and nothing else
catches it.
Skip rather than fail when hardware is absent (if (!HaveGpu()) GTEST_SKIP()),
so the suite stays runnable on a CPU-only box.
Register the test in tests/cpp/umbp/distributed/CMakeLists.txt with
add_test(NAME ... COMMAND ...), not gtest_discover_tests — the file
explains why (discovery bakes in build-time cmake paths and goes red when the
image's ctest runs it).
Building and running
The umbp C++ build needs protoc and grpc_cpp_plugin, which live in the mori
Docker image, not on the bare host:
docker exec <mori-container> bash -c "cd <repo> && \
cmake -S . -B build -G Ninja -DCMAKE_BUILD_TYPE=Release \
-DBUILD_UMBP=ON -DBUILD_TESTS=ON -DGPU_TARGETS=gfx950 \
-DCMAKE_CXX_COMPILER=/opt/rocm/lib/llvm/bin/clang++ \
-DCMAKE_HIP_COMPILER=/opt/rocm/lib/llvm/bin/clang++ && \
ninja -C build && cd build && ctest -E '^cco_' -LE integration"
Exclude ^cco_ — those are unrelated GPU collective tests and some hang without
a full fabric. -LE integration skips the tests needing real RDMA.
Reference
medium_backend.h— the interface, and a list of things deliberately not on it (with reasons, so they are not re-proposed)page_backend.h—PageMemorySource,HostPageMemorySource,PageBackendhbm_backend.{h,cpp}— the minimal paged medium (shape 1)ssd_backend.{h,cpp}— the staged, non-addressable medium (shape 2)doc/design-backend-agnostic-refactor.md— §2 the descriptor/pointer rule, §3 the three-component split, §8 what none of it fixes