About
The African DSI DataBank is a federated, hub-and-spoke platform for hosting, mirroring, curating, and governing access to Africa-origin biodiversity and agricultural Digital Sequence Information (DSI): genomes and transcriptomes, protein sequences and experimental proteomics, predicted protein structures, eDNA/metabarcoding and metagenomics, chemical (targeted) and metabolic (untargeted) profiles, and field/specimen and cellular/molecular imaging. Data flows upward from institution-owned Spoke nodes through National and Sub-regional Hubs to one Continental Hub, with full provenance preserved at every step, and can flow back down as real bytes — not just metadata — through a consent-gated transfer mechanism that always requires both sides to agree.
Capability maturity
A three-way split of what's actually running versus what's a placeholder. "Verified live" means actually run against a real external service, a real browser, or a real multi-service deployment — not just covered by the automated test suite.
| Capability | Status |
|---|---|
| Federated sync chain (Spoke → National → Sub-regional → Hub) | ✅ Verified live |
| Central JWT authentication, roles, node admission/approval | ✅ Verified live |
| Node<->Hub bidirectional data transfer (genomes, transcriptomes, proteins, eDNA/metagenomic/proteomics/metabolomics raw samples) | ✅ Verified live |
| Third-party node backend support (documented, versioned wire protocol + automatic capability handshake) | ✅ Verified live — verified against both a real stock node and a real, independently-written third-party node |
| Genome browser (JBrowse2) | ✅ Verified live — pixel-confirmed via headless-browser session |
| NCBI/ENA genome mirroring, GoaT assembly-quality enrichment, GBIF image mirroring | ✅ Verified live — against the real external APIs |
| Cellular/molecular imaging metadata lookup | ✅ Verified live — against the real external APIs |
| Production Docker deployment (Postgres/Redis/MinIO, non-root containers, rate limiting) | ✅ Verified live — built and run for real, including a live rate-limit flood test |
| Native (non-Docker) install — systemd units + .deb/.rpm packaging, Hub and node tiers | ✅ Verified live — verified on real Debian- and RHEL-family systemd hosts |
| API versioning (/api/v1) with real pagination on every Hub list endpoint | ✅ Verified live |
| Per-species data-completeness dashboard | ✅ Verified live |
| Faceted genome search (country, license, origin, assembly quality) | ✅ Verified live |
| Continental-Admin-controlled nav visibility (feature flags, both frontends) | ✅ Verified live |
| Structure viewer (Mol*) | 🟡 Implemented, unverified — not yet visually confirmed live — no browser available at build time |
| Bulk export (Darwin Core Archive / DCAT catalog) | ✅ Verified live — meta.xml/eml.xml structure verified against GBIF's own Darwin Core Text Guidelines |
| Compute-request execution (resource_type=blast) | ✅ Verified live — other resource types (phylogenetics, general_hpc) still have no execution path |
| DDBJ genome mirroring (file fetch specifically) | 🟡 Implemented, unverified — matches DDBJ's documented API shape, not confirmed against the live service |
| Third-party ABS (access-and-benefit-sharing) identity verification | 🟡 Implemented, unverified — built against a documented contract — no live third-party ABS service exists yet |
| Real DOI minting | 🟡 Implemented, unverified — scaffolded, inert until a DataCite prefix is registered |
| Transactional email (password reset, Contact Us) | 🟡 Implemented, unverified — two real implementations — a dedicated API (Resend) and a real SMTP mailbox — neither yet exercised against a live account/mailbox in this repo's own testing |
| Structure prediction (ESMFold) | 🔴 Pluggable stub — proves the job/versioning pipeline only |
| Protein domain annotation (Pfam/InterPro via EBI InterProScan5) | 🟡 Implemented, unverified — submit + status legs confirmed live; the result-fetch leg wasn't reached before a live job finished |
| eDNA classification (DADA2/QIIME2) | 🔴 Pluggable stub |
| Metagenomics profiling (Kraken2/HUMAnN) | 🔴 Pluggable stub |
| Proteomics / metabolomics sample analysis (mass-spec/NMR) | 🔴 Pluggable stub — the raw-sample-plus-derived-result sync/transfer pipeline itself is real and live-verified; only the analysis step is a stub, same as eDNA/metagenomics |
| Species/image ML classification, specimen-label OCR | 🔴 Pluggable stub |
| ENA submission — Study/Sample registration, Webin-CLI assembly submission, and accession polling (Phases 2-4) | 🟡 Implemented, unverified — written to ENA's own documented sandbox/Webin-CLI/Portal Search API contracts (a real, sha256-verified Webin-CLI JAR is bundled into the Hub's Docker image and native install) and covered by tests; not yet exercised against a live account — the user has not registered ENA Webin credentials yet. A settings-gated production-cutover switch (ENA_SUBMISSION_ENVIRONMENT) exists but defaults to sandbox everywhere. |
| Embargoed access tier — auto-release on a set date (genomes, Phase 1) | ✅ Verified live — a fourth access_tier alongside public/regional_researcher/restricted, with two submitter-chosen sub-levels during the embargo window and a real, live-verified scheduled release sweep plus a request-time lazy-release backstop. Genomes only this pass — 10 more citable models remain for future passes. |
Governance & policy
Governance (node admission, Continental override, dispute resolution), Terms of use & citation policy (data use and citation policy), and Preservation & sustainability (preservation and funding model).
