Node operator handbook
A practical overview for an institution considering running a Sub-regional Hub, National Hub, or Spoke — a node — in the African DSI DataBank federation. See Governance for how node admission and administrative decisions work. Requesting an operator account (Apply) and registering a brand-new node's existence (Register node) are two separate steps, but not independent ones — see "Getting a node running" below for the order that actually works.
Stay up to date: new features, action needed and maintenance are announced in Updates for node operators, which every node also shows on its own Updates page.
This page is an overview. The full handbook, with every setting and command, comes with the node software itself: docs/NODE_OPERATOR_HANDBOOK.md in your node's files, updated with each release.
What running a node means
Every tier — Sub-regional Hub, National Hub, Spoke — runs the same underlying software, configured differently. Your node holds your own institution's data as the source of truth, and decides for itself whether and what to sync upward to the rest of the federation.
Hardware & bandwidth
There is no fixed hardware requirement — the node software itself is small. What actually drives sizing is your own data:
- A Spoke capturing genome assemblies, proteins, or raw eDNA/metagenomic reads needs disk space proportional to what it captures — plan around your own expected data volume, not a platform-wide number.
- A National or Sub-regional Hub additionally aggregates what its own children sync upward — metadata and derived results are far smaller than the raw files a Spoke may hold locally and never sync.
- A single small server is enough to run the software itself; size disk and bandwidth to your own data, not the compute.
Getting a node running
- Request an operator account first — see Apply. You'll state whether this is for a node that already exists (type its known node ID) or a brand-new one (the system generates a real, unique ID for you and shows it back once your request is approved — copy it down, you'll need it in step 2). Approval is a manual, human decision by the Continental Infrastructure Admin, and also mints a single-use applicant ID, emailed to you.
- Register the node's existence — brand-new nodes only. If you applied against an already-existing node, skip this step entirely; your account is already active and scoped to it the moment step 1 is approved. For a genuinely new node, use your applicant ID (plus the exact tier and node ID from step 1 — they must match) at Register node: a new Sub-regional Hub applies directly to the Continental Hub; a new National Hub or Spoke applies to its intended parent instead. This starts pending until that parent approves it too.
- Deploy the software. Three genuinely supported options: a Docker-based deployment, a native (non-Docker) install via a script or a
.deb/.rpmpackage, or plain local processes for development only. Pick whichever fits your own comfort level and infrastructure. Your node can have its own host name or live under a path on an existing site (for examplehttps://example.org/AfricanDSIDataBank/, set withNODE_BASE_PATH). - Configure it in
node.env, never by editing the code. Updates replace the code folders, so every adjustment an operator needs is a setting innode.env, or, for Docker additions such as an extra folder to mount, your own compose file kept outside thedeploy/folder. - Verify it's healthy — a real health check confirms your node's database and storage are genuinely reachable, not just that the process started.
Getting the software, and verifying it's genuine
This project's own source code is private, and you don't need git access to it — every release (whichever of the three install options above you pick) is a set of plain, downloadable files, gated to logged-in, approved node operators. Two independent checks are made against every download before it's ever trusted: its manifest's own GPG signature is verified against this project's published release-signing key, and the file's own SHA-256 checksum is then verified against that now-trusted manifest. Either check failing means the file is rejected outright.
The release-signing key's fingerprint (ed25519, sign-only, never expires) — published here specifically so you have somewhere independent of the download itself to cross-check against, the standard "trust on first use" step this signing scheme depends on:
D7B5 B8A7 58C7 2D04 B778 1F36 5D29 8472 953F 3E09
A copy-pasteable first-install example (only the password prompt is ever typed — it's read straight into a shell variable, never echoed or logged):
HUB_URL="https://hub.africandsidatabank.africa"
# 1. Find the current release's id
curl -sS "$HUB_URL/api/v1/node-releases/latest"
# -> {"id": 3, "version": "0.1.2", ...} - note the "id"
# 2. Log in with your own central-login account
read -p "Username: " NODE_USER
read -sp "Password: " NODE_PASS; echo
JWT=$(curl -sS -X POST "$HUB_URL/api/v1/accounts/login" \
-H "Content-Type: application/json" \
-d "{\"username\": \"$NODE_USER\", \"password\": \"$NODE_PASS\"}" \
| python3 -c "import sys,json; print(json.load(sys.stdin)['access_token'])")
# 3. Download the manifest, its signature, and the gated tarball (use the
# id from step 1; swap "tarball" for "deb"/"rpm" if you'd rather use a
# package)
RELEASE_ID=3
curl -sS -o SHA256SUMS.txt "$HUB_URL/api/v1/node-releases/$RELEASE_ID/checksums-manifest"
curl -sS -o SHA256SUMS.txt.asc "$HUB_URL/api/v1/node-releases/$RELEASE_ID/checksums-manifest.asc"
curl -sS -H "Authorization: Bearer $JWT" -o node-release.tar.gz \
"$HUB_URL/api/v1/node-releases/$RELEASE_ID/download/tarball"
# 4. Extract ONLY deploy/ first, then cross-check the fingerprint it
# reports against the value published above, before trusting anything
# further
tar xzf node-release.tar.gz ./deploy
gpg --show-keys --with-fingerprint ./deploy/RELEASE_SIGNING_PUBLIC_KEY.asc
# 5. Once the fingerprint genuinely matches, verify the whole download -
# it exits non-zero on any failure, and only prints what to run next
# (extract the rest, then install) once both checks genuinely pass
./deploy/verify-node-release.sh node-release.tar.gzPrefer not to run curl commands by hand? Once you're approved and logged in, your own node's "Tier status" page (see "Upgrading" below) can download and verify a release for you the same way — the manual walkthrough above is the one thing worth doing yourself at least once, so you've seen the fingerprint check with your own eyes before trusting it to run unattended.
If the Hub can't reach your node: pushing your changes
Normally your parent node and the Continental Hub each fetch your node's changes on a schedule. If they can't reach your node (often because an institutional firewall blocks unfamiliar servers), your node can send its changes instead: to the Continental Hub, to its parent, or to both. You choose.
- Ask each side you want to push to to allow it: the Continental Admin for the Hub, and your parent's operator for your parent.
- On your node's Tier status page, click Get token for each side. Your node fetches its own push token with your login, so nothing needs copying or sending.
- Tick where to push. Your node then pushes on its regular sync schedule, and the same page has a Push now button and shows how each push went.
Pushed changes go through exactly the same checks as fetched ones, and push only replaces sync: transfers and compute requests still need the Hub to reach your node.
Using files already on your server
In addition to uploading files through the browser, your node can use data files that already sit on its server. Your administrator lists the folders it may use (NODE_DATA_ROOTS in node.env). Then every upload form offers "A file already on this server", with a folder browser, and the node's Files on this server page imports many records at once from a CSV, for any data type.
For each file you choose either to copy it into the node, like an upload, or to serve it in place, with no copy. A file served in place must stay where it is, unchanged: the node checks it regularly and stops serving it if it has moved or changed, until you re-check or re-register it. Records made either way are ordinary records: they sync, can be transferred and are citable like any other.
Taking part in federated training
From 0.5.0 your node can help train a model across nodes without its data leaving your server. The Continental Admin starts a training run and invites nodes; you accept or decline each run on your node's Federated training page, which shows how many of your records it would use. Your node then trains the run's model on its own records (public ones by default, never embargoed ones, never records relayed from other nodes) and sends the Hub only the model's weights and a few summary numbers. It connects out to the Hub itself, so it works behind a firewall.
Training runs in containers started by your node's container runner, with no network and fixed limits, using only images pinned to an exact version by the Hub (from 0.5.1 usually served by the Hub itself, only to nodes taking part in the run). It's off until you switch it on (NODE_FEDERATED_TRAINING, plus the runner); the full handbook explains how, including GPUs. The finished model is published on the Hub as a citable result listing which nodes contributed.
Bringing your own node backend
You don't have to run this platform's own reference software. The federation's node protocol is a documented, versioned contract — if your institution already runs its own bioinformatics infrastructure, you can implement that contract directly against your own systems instead of standing up a second stack. The minimum to be a working node is small; everything beyond that is optional and degrades gracefully if you don't implement it — that one feature simply doesn't work for your node, nothing else is affected. The Continental Hub will label your node by its own reported software name rather than assuming it runs the reference implementation — purely informational, never a gate on anything.
Backup
Your node's database (metadata) and object storage (actual file bytes) are two separate things and both need regular backups — losing either loses real data your node is the source of truth for. Content checksums, computed when a file reaches the Hub, are one way to later verify a restored copy matches what was originally captured. Files your node serves in place from its own folders aren't in its storage, so back those folders up as well.
Upgrading
Database schema changes apply automatically on startup — in practice, get a newer version and restart the service, and the schema updates itself before the software starts serving requests. Take a backup before any version upgrade regardless.
A banner tells you when a newer version is available; it never applies anything on its own, and it turns amber for a security release and red once an update has waited too long. Your node's Tier status page then offers Update now, Update in 3 hours or Defer. You can also turn on automatic updates there (off by default), and your node then checks about every 15 minutes and installs a new version itself, after verifying it.
Before updating, the updater checks that nobody has edited the node's code. If someone has, it stops and lists the changed files rather than overwrite them. It backs up the current code before every update and never touches your node.env. Packages installed with .deb/.rpm are upgraded by installing the newer package.
Recognition for what your node contributes
Every DSI record your node syncs upward gets a real, durable citable id the moment it reaches the Hub, with a one-click citation string and BibTeX export on its own citation page — see Terms of Use's citation policy. Your node's own real contribution counts (records, views, downloads) are also visible, publicly and by name, on the Hub's Participation dashboard.
Getting help
- A security vulnerability — see Security for how to report it privately.
- A dispute over an administrative decision — see Governance.
- Everything else — reach the Continental Hub operator through Contact Us (also on the homepage).
See also: Governance, API guide.
