API docs

The African DSI DataBank is a federated catalog of African-origin biodiversity/agriculture Digital Sequence Information — genomes, proteins, predicted structures, eDNA/metagenomic classifications, and field/specimen/cellular imaging metadata, aggregated from Sub-regional, National, and Spoke nodes across the continent. Every endpoint below is real and live; the interactive, always-current reference is the OpenAPI schema itself: /api/docs ↗.

Access tiers

Every record carries one of three access tiers, enforced before any byte is served — see Terms of use for the full explanation.

Authentication

Public reads need no token. Write endpoints and anything beyond `public`-tier data require a bearer token from POST /api/v1/accounts/login:

curl "https://<hub-domain>/api/v1/accounts/login" \
  -X POST -H "Content-Type: application/json" \
  -d '{"username":"you","password":"..."}'

The response's access_token goes in an Authorization: Bearer <token> header on subsequent requests.

Common endpoints

Search the species catalog

curl "https://<hub-domain>/api/v1/catalog/species?q=Loxodonta"

List genome sequence records

curl "https://<hub-domain>/api/v1/catalog/sequences"

Download a genome's assembly file

curl "https://<hub-domain>/api/v1/catalog/sequences/{id}/file"

List proteins

curl "https://<hub-domain>/api/v1/proteins"

List predicted structures

curl "https://<hub-domain>/api/v1/structures"

eDNA / metagenomics per-site dashboards

curl "https://<hub-domain>/api/v1/edna/sites"
curl "https://<hub-domain>/api/v1/metagenomics/sites"

Resolve a citable id

curl "https://<hub-domain>/api/v1/identifiers/AFDSI-SEQ-1042"

Public federation contribution evidence

curl "https://<hub-domain>/api/v1/federation/contributions"

Bulk & machine-readable export

Two public, unauthenticated feeds for external harvesting/discovery tooling — neither needs an account, and both are hard-scoped to public-tier records only.

Darwin Core Archive (all public genomes)

A standard Darwin Core Archive ZIP (meta.xml + an occurrence-core data table + eml.xml) — the format GBIF's own IPT/harvesting pipeline consumes. Scoped to genomes only, since Darwin Core's Occurrence core is specimen/collection- event-shaped and only genomes carry that data in this catalog.

curl "https://<hub-domain>/api/v1/export/darwin-core.zip"

DCAT catalog description

A JSON-LD dcat:Catalog document listing every dataset this Hub holds (all 7 citable entity types) with record counts and links to each dataset's own list endpoint — the machine-readable "what does this Hub have" document a national or regional open-data portal can index.

curl "https://<hub-domain>/api/v1/export/dcat"

A short Python example

import requests

BASE = "https://<hub-domain>"

# Search the species catalog — no auth needed for public data.
# List endpoints are paginated: {"items": [...], "count": N}.
response = requests.get(f"{BASE}/api/v1/catalog/species", params={"q": "Loxodonta"})
response.raise_for_status()
for species in response.json()["items"]:
    print(species["scientific_name"], species["common_name"])

# Anything beyond public-tier data needs a bearer token first.
login = requests.post(
    f"{BASE}/api/v1/accounts/login",
    json={"username": "you", "password": "..."},
).json()
headers = {"Authorization": f"Bearer {login['access_token']}"}
requests.get(f"{BASE}/api/v1/access/requests", headers=headers).raise_for_status()

A full worked example: search → request access → download

The short snippet above shows individual calls in isolation. This complete, runnable script (also in the repository at docs/examples/researcher_workflow_example.py) composes them into the real end-to-end flow a researcher actually needs: search the catalog, find a genome record, and either download it directly (public-tier) or submit a real access request and check on its status (anything else). Two details worth knowing before reading it: submitting an access request needs no login at all — it's free-text, verified against who is logged in later when a download is actually attempted, not at submission time — and, on a deployment with a real CAPTCHA key configured, that same submission needs a token a headless script can't obtain on its own (use the real form in that case). Run it with python researcher_workflow_example.py "Species name" — only dependency is pip install requests.

import os
import sys

import requests

BASE_URL = os.environ.get("HUB_BASE_URL", "https://hub.africandsidatabank.africa")
USERNAME = os.environ.get("HUB_USERNAME", "")
PASSWORD = os.environ.get("HUB_PASSWORD", "")


def log(message):
    print(f"\n>>> {message}")


# Step 1: search the species catalog — no auth needed, ever. List
# endpoints are paginated: {"items": [...], "count": N}.
def find_species(scientific_name):
    log(f"Searching the species catalog for '{scientific_name}'...")
    response = requests.get(
        f"{BASE_URL}/api/v1/catalog/species", params={"q": scientific_name}, timeout=30
    )
    response.raise_for_status()
    matches = response.json()["items"]
    if not matches:
        raise SystemExit(f"No species matched '{scientific_name}'.")
    species = matches[0]
    print(f"Found: {species['scientific_name']} (species id {species['id']})")
    return species


# Step 2: find a genome record for that species — also public.
def find_genome_record(species_id):
    log(f"Looking up sequence records for species id {species_id}...")
    response = requests.get(
        f"{BASE_URL}/api/v1/catalog/sequences",
        params={"species_id": species_id, "sequence_type": "genome"},
        timeout=30,
    )
    response.raise_for_status()
    records = response.json()["items"]
    if not records:
        print("No genome-type sequence record exists for this species yet.")
        return None
    record = records[0]
    print(
        f"Found sequence record {record['id']} "
        f"(access_tier={record['access_tier']!r}, citable_id={record['citable_id']!r})"
    )
    return record


# Step 3a: the public-tier path — download directly, no auth.
def download_genome(sequence_record_id, out_path, headers=None):
    log(f"Downloading assembly file for sequence record {sequence_record_id}...")
    response = requests.get(
        f"{BASE_URL}/api/v1/catalog/sequences/{sequence_record_id}/file",
        headers=headers or {},
        timeout=120,
    )
    response.raise_for_status()
    with open(out_path, "wb") as handle:
        handle.write(response.content)
    print(f"Saved {len(response.content):,} bytes to {out_path}")


# Step 3b: the gated path — submitting a request needs no bearer token at
# all; checking status and any actual download does. require_access only
# ever grants a download to the exact account named as "requester" on an
# already-approved request (or a Continental Admin).
def request_access(sequence_record_id, requester):
    log(f"Submitting an access request for sequence record {sequence_record_id}...")
    response = requests.post(
        f"{BASE_URL}/api/v1/access/requests",
        json={
            "sequence_record_id": sequence_record_id,
            "requester": requester,
            "purpose": "Worked-example script: evaluating this genome for a comparative study.",
        },
        timeout=30,
    )
    response.raise_for_status()
    request = response.json()
    print(
        f"Access request #{request['id']} submitted — status: {request['decision']!r}. "
        "A Continental Admin needs to approve this before the file can be downloaded."
    )
    return request


def login(username, password):
    log(f"Logging in as {username!r}...")
    response = requests.post(
        f"{BASE_URL}/api/v1/accounts/login",
        json={"username": username, "password": password},
        timeout=30,
    )
    response.raise_for_status()
    return response.json()["access_token"]


def check_access_request_status(request_id, headers):
    log(f"Checking status of access request #{request_id}...")
    response = requests.get(
        f"{BASE_URL}/api/v1/access/requests/{request_id}", headers=headers, timeout=30
    )
    response.raise_for_status()
    decision = response.json()["decision"]
    print(f"Access request #{request_id} is currently: {decision!r}")
    return decision


# GET /requests, logged in, already returns only your own requests — this
# keeps the script safe to re-run instead of piling up duplicate requests.
def find_own_request_for_record(sequence_record_id, headers):
    response = requests.get(f"{BASE_URL}/api/v1/access/requests", headers=headers, timeout=30)
    response.raise_for_status()
    for existing in response.json()["items"]:
        if existing["sequence_record_id"] == sequence_record_id:
            return existing
    return None


def main():
    scientific_name = sys.argv[1] if len(sys.argv) > 1 else "Loxodonta africana"

    species = find_species(scientific_name)
    record = find_genome_record(species["id"])
    if record is None:
        return

    out_path = f"{record['citable_id'] or record['id']}.fasta"

    if record["access_tier"] == "public":
        download_genome(record["id"], out_path)
        return

    print(
        f"\nThis record is '{record['access_tier']}'-tier, not public — a download needs an "
        "approved access request first."
    )
    if not USERNAME:
        print(
            "No HUB_USERNAME set — register an account (see /apply), then re-run this script "
            "with HUB_USERNAME=<your username> to submit a real access request under your own "
            "name (no login needed for this step)."
        )
        return

    if not PASSWORD:
        request_access(record["id"], requester=USERNAME)
        print(
            "\nSet HUB_PASSWORD too, once a Continental Admin has approved this request, to "
            "actually log in and download — re-run this exact command with both set."
        )
        return

    token = login(USERNAME, PASSWORD)
    headers = {"Authorization": f"Bearer {token}"}

    request = find_own_request_for_record(record["id"], headers)
    if request is None:
        request = request_access(record["id"], requester=USERNAME)

    status = request["decision"]
    if status != "approved":
        status = check_access_request_status(request["id"], headers)

    if status == "approved":
        download_genome(record["id"], out_path, headers=headers)
    else:
        print(
            "\nNot approved yet. Re-run this script once a Continental Admin approves it — "
            "the download will then proceed automatically."
        )


if __name__ == "__main__":
    main()

See also