Python SDK

Push, pull, search, and manage artifacts from Python with the official artefaktum package

Install

pip install artefaktum

Python 3.10+. The only dependency is httpx. For the CLI, see CLI.

Configure

Resolved in order: constructor argument, then environment variable, then default.

Argument Environment variable Default
api_key ARTEFAKTUM_API_KEY (required)
base_url ARTEFAKTUM_BASE_URL https://api.artefaktum.dev
project ARTEFAKTUM_PROJECT "default"
client = Artefaktum(api_key="ak_...", base_url="http://localhost:3000", project="default", timeout=30.0)

API keys are created in the console. MissingApiKey (a ValueError) is raised locally, before any request, if no key is found anywhere.

project is a slug or a UUID. A slug is resolved to an ID through GET /v1/projects once, then cached. push, search, list, resolve, and a few others also take a per-call project= override.

Five-minute example

from artefaktum import Artefaktum

client = Artefaktum()  # ARTEFAKTUM_API_KEY, project "default"
a = client.artifacts.push("report.pdf", title="Q3 churn", tags=["churn"])
hits = client.artifacts.search("q3 churn")
client.artifacts.pull(hits.items[0].artifact.id, "out/")  # sha256 verified

Pushing

push hashes, reserves, uploads, and (by default) waits for the artifact to be ready:

  1. Hashes and sizes the source locally, streamed; guesses a content type from the filename unless one is given.
  2. Reserves an upload (create_upload), returning a signed PUT URL.
  3. Uploads the bytes directly to object storage, over a client with no Authorization header, so the API key never reaches the storage host.
  4. Marks the version complete (complete_upload) with the digest and size.
  5. By default, polls every 0.5s until ready, raising ProcessingFailed or ProcessingTimeout. wait=False returns right after step 4, with status processing or ready.
client.artifacts.push(
    "report.pdf",
    title="Q3 churn",
    description="Churn breakdown by segment",
    tags=["churn", "q3"],
    external_key="reports/q3-churn",
)
Argument Default Meaning
source required path (streamed) or bytes (needs filename)
title required artifact title
description ""
tags () sequence of tags
metadata None mapping of custom fields
external_key None key for resolve / get_by_external_key
expires_at None datetime after which the artifact expires
run None attach to a run ID
infer_lineage True let the server infer relations
content_type None guessed from filename if omitted
filename None required for bytes; else the path’s name
wait True poll until ready
timeout 30.0 seconds to wait
project None override the default project

Field sizes are limited; see Limits.

create_version(artifact_id, source, ...) runs the same flow for a new version of an existing artifact: run, infer_lineage, content_type, filename, wait, timeout, but no title/tags/metadata, since those belong to the artifact, not the version.

Pulling

path = client.artifacts.pull(artifact_id, "out/")

pull(artifact_id, dest, *, verify=True) streams the latest version to dest and returns the final Path. dest is a directory when it already is one, or when it is a string ending in /, which pull creates. A pathlib.Path cannot express that intent, since Path("out/") is just Path("out"). Into a directory, the file is named after the version’s stored filename, reduced to a basename, so an uploader cannot steer the write outside it.

The download runs on a client with no Authorization header. verify=True (the default) checks the sha256 digest and raises IntegrityError on a mismatch, deleting the partial file; a version the server recorded no digest for is written unverified regardless. verify=False skips the check.

Finding

hits = client.artifacts.search("q3 churn", mode="hybrid", tags_all=["churn"])
page = client.artifacts.list(tag="churn", limit=50)
one = client.artifacts.get(artifact_id)
one = client.artifacts.get_by_external_key("reports/q3-churn")

search returns a SearchPage of SearchHit (.artifact, .score, .match_mode). Key arguments: query, mode (hybrid default, or text, semantic, exact), limit (20), cursor, content_types, tags_all, status (default ("ready",)), external_key, created_after/created_before, exclude_superseded (True by default), project.

list returns a Page[Artifact], with the same filters as search minus query/mode, plus expires_before and default limit=50. iter_all(...) takes the same filters and is a generator that follows next_cursor page by page, yielding every matching Artifact; on the async client, iter_all is an async generator instead.

get(artifact_id) and get_by_external_key(key, *, project=None) each return a single Artifact, raising NotFound if there is none.

Resolving and caching by external key

resolve is get-or-create for an external key, so a run avoids re-uploading something another run already produced:

resolution = client.artifacts.resolve(
    "reports/q3-churn", filename="report.pdf", content_type="application/pdf",
    size_bytes=len(data), title="Q3 churn", max_age=timedelta(days=30),
)
if resolution.status == "create":
    artifact = client.artifacts.fulfil(resolution, "report.pdf")
else:
    artifact = resolution.artifact  # status == "hit"

max_age (int, float, or timedelta seconds, sent as max_age_seconds) treats an older artifact as stale, so resolve returns status="create" instead of "hit".

Resolution.status: hit (.artifact set, nothing uploaded), create (.reservation/.upload set; pass the resolution and bytes to fulfil(resolution, source, *, filename=None, wait=True, timeout=30.0), which runs push’s upload flow), or pending (another writer holds the reservation; .retry_after_seconds hints the wait). fulfil raises ValueError for any other status.

Versions and relations

versions = client.artifacts.versions(artifact_id)
relation = client.artifacts.add_relation(artifact_id, other_id, "derived_from")
relations = client.artifacts.relations(artifact_id)
client.artifacts.update(artifact_id, title="New title", tags=["a", "b"])
client.artifacts.delete(artifact_id)

versions(artifact_id) returns every Version, newest last. add_relation(artifact_id, to_artifact_id, relation_type, *, to_version_id=None, metadata=None) requires relation_type to be derived_from, supersedes, attachment_of, generated_by, or related_to, else ValueError. relations(artifact_id) lists relations both ways. update(...) patches title, description, tags, metadata, expires_at, clear_expires_at without touching the file. delete(artifact_id) removes an artifact.

Async client

AsyncArtefaktum mirrors every method on Artefaktum as a coroutine; iter_all becomes an async generator:

import asyncio
from artefaktum import AsyncArtefaktum


async def main():
    async with AsyncArtefaktum() as client:
        a = await client.artifacts.push("report.pdf", title="Q3 churn")
        async for artifact in client.artifacts.iter_all(tag="churn"):
            print(artifact.id)


asyncio.run(main())

Errors

Every failure is an ArtefaktumError(message, code, status, request_id) or a subclass:

from artefaktum import Artefaktum, NotFound, ArtefaktumError

with Artefaktum() as client:
    try:
        client.artifacts.get("01a0be89-1cde-74c1-ab17-8814cc9d9141")
    except NotFound as e:
        print(e.code, e.request_id)  # "artifact_not_found", "req_..."
    except ArtefaktumError as e:
        print(str(e))  # "<code>: <message> (request_id=<id>)"
Exception Code(s) Status When
NotFound artifact_not_found, run_not_found, not_found 404 no such artifact or run
Unauthorized unauthorized 401 bad API key
Forbidden insufficient_scope 403 key lacks scope
QuotaExceeded quota_exceeded 413, 429 413: file/storage too large; 429: calls spent
RequestTooLarge request_too_large 413 the JSON request is over 256 KB
Conflict artifact_not_ready, external_key_conflict, idempotency_conflict, run_sealed 409 server-side conflict
ValidationFailed invalid_request 400 malformed request
UploadError upload_expired, object_verification_failed 410, 422 upload window passed, or verification failed
ServiceUnavailable embedding_unavailable 503 embeddings unavailable
ProcessingFailed processing_failed n/a wait saw status failed; carries .artifact
ProcessingTimeout processing_timeout n/a wait exceeded timeout; carries .artifact
IntegrityError integrity_error n/a pull sha256 mismatch; .expected/.actual
StorageError storage_error n/a storage host failed; carries .host
MissingApiKey n/a n/a raised locally, before any request

A 429 QuotaExceeded blocks writes and search until the plan resets; reads, downloads, and deletes keep working. client.quota() reports the plan, limits, and usage, and is never blocked:

q = client.quota()
print(q.plan, q.storage.used_bytes, "/", q.storage.limit_bytes, q.calls.resets_at)

Retries and timeouts

Reads retry automatically on HTTP 429, 502, 503, or 504, and on connection errors, up to 3 attempts, waiting 0.5s then 1s (or the server’s Retry-After, capped at 10s): get, get_by_external_key, list / iter_all, search, download_url, whoami, projects.list, usage.get. search is a POST but has no side effects, so it retries too. A 429 coded quota_exceeded is the exception, raised immediately since the allowance returns only at quota().calls.resets_at.

Writes (push, update, add_relation, delete, keys.create, and the rest) are never retried automatically, so a client never double-submits one; neither is the storage upload or download.

The timeout constructor argument (default 30.0) sets the HTTP timeout on both the API and storage clients. The separate timeout= on push, create_version, and fulfil controls only how long the ready-poll waits, independent of the HTTP timeout.

This page as Markdown: /docs/python.md