# Python SDK

> Push, pull, search, and manage artifacts from Python with the official artefaktum package

## Install

```sh
pip install artefaktum
```

Python 3.10+. The only dependency is `httpx`. For the CLI, see [CLI](/docs/cli/).

## Configure

Resolved in order: constructor argument, then environment variable, then default.

| Argument | Environment variable | Default |
|---|---|---|
| `api_key` | `ARTEFAKTUM_API_KEY` | *(required)* |
| `base_url` | `ARTEFAKTUM_BASE_URL` | `https://api.artefaktum.dev` |
| `project` | `ARTEFAKTUM_PROJECT` | `"default"` |

```python
client = Artefaktum(api_key="ak_...", base_url="http://localhost:3000", project="default", timeout=30.0)
```

API keys are created in the [console](/console/). `MissingApiKey` (a `ValueError`) is raised locally, before any request, if no key is found anywhere.

`project` is a slug or a UUID. A slug is resolved to an ID through `GET /v1/projects` once, then cached. `push`, `search`, `list`, `resolve`, and a few others also take a per-call `project=` override.

## Five-minute example

```python
from artefaktum import Artefaktum

client = Artefaktum()  # ARTEFAKTUM_API_KEY, project "default"
a = client.artifacts.push("report.pdf", title="Q3 churn", tags=["churn"])
hits = client.artifacts.search("q3 churn")
client.artifacts.pull(hits.items[0].artifact.id, "out/")  # sha256 verified
```

## Pushing

`push` hashes, reserves, uploads, and (by default) waits for the artifact to be ready:

1. Hashes and sizes the source locally, streamed; guesses a content type from the filename unless one is given.
2. Reserves an upload (`create_upload`), returning a signed PUT URL.
3. Uploads the bytes directly to object storage, over a client with no `Authorization` header, so the API key never reaches the storage host.
4. Marks the version complete (`complete_upload`) with the digest and size.
5. By default, polls every 0.5s until `ready`, raising `ProcessingFailed` or `ProcessingTimeout`. `wait=False` returns right after step 4, with status `processing` or `ready`.

```python
client.artifacts.push(
    "report.pdf",
    title="Q3 churn",
    description="Churn breakdown by segment",
    tags=["churn", "q3"],
    external_key="reports/q3-churn",
)
```

| Argument | Default | Meaning |
|---|---|---|
| `source` | required | path (streamed) or `bytes` (needs `filename`) |
| `title` | required | artifact title |
| `description` | `""` | |
| `tags` | `()` | sequence of tags |
| `metadata` | `None` | mapping of custom fields |
| `external_key` | `None` | key for `resolve` / `get_by_external_key` |
| `expires_at` | `None` | `datetime` after which the artifact expires |
| `run` | `None` | attach to a run ID |
| `infer_lineage` | `True` | let the server infer relations |
| `content_type` | `None` | guessed from `filename` if omitted |
| `filename` | `None` | required for `bytes`; else the path's name |
| `wait` | `True` | poll until `ready` |
| `timeout` | `30.0` | seconds to wait |
| `project` | `None` | override the default project |

Field sizes are limited; see [Limits](/docs/rest/#limits).

`create_version(artifact_id, source, ...)` runs the same flow for a new version of an existing artifact: `run`, `infer_lineage`, `content_type`, `filename`, `wait`, `timeout`, but no title/tags/metadata, since those belong to the artifact, not the version.

## Pulling

```python
path = client.artifacts.pull(artifact_id, "out/")
```

`pull(artifact_id, dest, *, verify=True)` streams the latest version to `dest` and returns the final `Path`. `dest` is a directory when it already is one, or when it is a *string* ending in `/`, which `pull` creates. A `pathlib.Path` cannot express that intent, since `Path("out/")` is just `Path("out")`. Into a directory, the file is named after the version's stored filename, reduced to a basename, so an uploader cannot steer the write outside it.

The download runs on a client with no `Authorization` header. `verify=True` (the default) checks the sha256 digest and raises `IntegrityError` on a mismatch, deleting the partial file; a version the server recorded no digest for is written unverified regardless. `verify=False` skips the check.

## Finding

```python
hits = client.artifacts.search("q3 churn", mode="hybrid", tags_all=["churn"])
page = client.artifacts.list(tag="churn", limit=50)
one = client.artifacts.get(artifact_id)
one = client.artifacts.get_by_external_key("reports/q3-churn")
```

`search` returns a `SearchPage` of `SearchHit` (`.artifact`, `.score`, `.match_mode`). Key arguments: `query`, `mode` (`hybrid` default, or `text`, `semantic`, `exact`), `limit` (20), `cursor`, `content_types`, `tags_all`, `status` (default `("ready",)`), `external_key`, `created_after`/`created_before`, `exclude_superseded` (`True` by default), `project`.

`list` returns a `Page[Artifact]`, with the same filters as `search` minus `query`/`mode`, plus `expires_before` and default `limit=50`. `iter_all(...)` takes the same filters and is a generator that follows `next_cursor` page by page, yielding every matching `Artifact`; on the async client, `iter_all` is an async generator instead.

`get(artifact_id)` and `get_by_external_key(key, *, project=None)` each return a single `Artifact`, raising `NotFound` if there is none.

## Resolving and caching by external key

`resolve` is get-or-create for an external key, so a run avoids re-uploading something another run already produced:

```python
resolution = client.artifacts.resolve(
    "reports/q3-churn", filename="report.pdf", content_type="application/pdf",
    size_bytes=len(data), title="Q3 churn", max_age=timedelta(days=30),
)
if resolution.status == "create":
    artifact = client.artifacts.fulfil(resolution, "report.pdf")
else:
    artifact = resolution.artifact  # status == "hit"
```

`max_age` (`int`, `float`, or `timedelta` seconds, sent as `max_age_seconds`) treats an older artifact as stale, so `resolve` returns `status="create"` instead of `"hit"`.

`Resolution.status`: `hit` (`.artifact` set, nothing uploaded), `create` (`.reservation`/`.upload` set; pass the resolution and bytes to `fulfil(resolution, source, *, filename=None, wait=True, timeout=30.0)`, which runs `push`'s upload flow), or `pending` (another writer holds the reservation; `.retry_after_seconds` hints the wait). `fulfil` raises `ValueError` for any other status.

## Versions and relations

```python
versions = client.artifacts.versions(artifact_id)
relation = client.artifacts.add_relation(artifact_id, other_id, "derived_from")
relations = client.artifacts.relations(artifact_id)
client.artifacts.update(artifact_id, title="New title", tags=["a", "b"])
client.artifacts.delete(artifact_id)
```

`versions(artifact_id)` returns every `Version`, newest last. `add_relation(artifact_id, to_artifact_id, relation_type, *, to_version_id=None, metadata=None)` requires `relation_type` to be `derived_from`, `supersedes`, `attachment_of`, `generated_by`, or `related_to`, else `ValueError`. `relations(artifact_id)` lists relations both ways. `update(...)` patches `title`, `description`, `tags`, `metadata`, `expires_at`, `clear_expires_at` without touching the file. `delete(artifact_id)` removes an artifact.

## Async client

`AsyncArtefaktum` mirrors every method on `Artefaktum` as a coroutine; `iter_all` becomes an async generator:

```python
import asyncio
from artefaktum import AsyncArtefaktum


async def main():
    async with AsyncArtefaktum() as client:
        a = await client.artifacts.push("report.pdf", title="Q3 churn")
        async for artifact in client.artifacts.iter_all(tag="churn"):
            print(artifact.id)


asyncio.run(main())
```

## Errors

Every failure is an `ArtefaktumError(message, code, status, request_id)` or a subclass:

```python
from artefaktum import Artefaktum, NotFound, ArtefaktumError

with Artefaktum() as client:
    try:
        client.artifacts.get("01a0be89-1cde-74c1-ab17-8814cc9d9141")
    except NotFound as e:
        print(e.code, e.request_id)  # "artifact_not_found", "req_..."
    except ArtefaktumError as e:
        print(str(e))  # "<code>: <message> (request_id=<id>)"
```

| Exception | Code(s) | Status | When |
|---|---|---|---|
| `NotFound` | `artifact_not_found`, `run_not_found`, `not_found` | 404 | no such artifact or run |
| `Unauthorized` | `unauthorized` | 401 | bad API key |
| `Forbidden` | `insufficient_scope` | 403 | key lacks scope |
| `QuotaExceeded` | `quota_exceeded` | 413, 429 | 413: file/storage too large; 429: calls spent |
| `RequestTooLarge` | `request_too_large` | 413 | the JSON request is over 256 KB |
| `Conflict` | `artifact_not_ready`, `external_key_conflict`, `idempotency_conflict`, `run_sealed` | 409 | server-side conflict |
| `ValidationFailed` | `invalid_request` | 400 | malformed request |
| `UploadError` | `upload_expired`, `object_verification_failed` | 410, 422 | upload window passed, or verification failed |
| `ServiceUnavailable` | `embedding_unavailable` | 503 | embeddings unavailable |
| `ProcessingFailed` | `processing_failed` | n/a | wait saw status `failed`; carries `.artifact` |
| `ProcessingTimeout` | `processing_timeout` | n/a | wait exceeded `timeout`; carries `.artifact` |
| `IntegrityError` | `integrity_error` | n/a | `pull` sha256 mismatch; `.expected`/`.actual` |
| `StorageError` | `storage_error` | n/a | storage host failed; carries `.host` |
| `MissingApiKey` | n/a | n/a | raised locally, before any request |

A 429 `QuotaExceeded` blocks writes and `search` until the plan resets; reads, downloads, and deletes keep working. `client.quota()` reports the plan, limits, and usage, and is never blocked:

```python
q = client.quota()
print(q.plan, q.storage.used_bytes, "/", q.storage.limit_bytes, q.calls.resets_at)
```

## Retries and timeouts

Reads retry automatically on HTTP 429, 502, 503, or 504, and on connection errors, up to 3 attempts, waiting 0.5s then 1s (or the server's `Retry-After`, capped at 10s): `get`, `get_by_external_key`, `list` / `iter_all`, `search`, `download_url`, `whoami`, `projects.list`, `usage.get`. `search` is a POST but has no side effects, so it retries too. A 429 coded `quota_exceeded` is the exception, raised immediately since the allowance returns only at `quota().calls.resets_at`.

Writes (`push`, `update`, `add_relation`, `delete`, `keys.create`, and the rest) are never retried automatically, so a client never double-submits one; neither is the storage upload or download.

The `timeout` constructor argument (default `30.0`) sets the HTTP timeout on both the API and storage clients. The separate `timeout=` on `push`, `create_version`, and `fulfil` controls only how long the ready-poll waits, independent of the HTTP timeout.
