Skip to main content

Answer from images and audio

By the end of this tutorial you will have an agent that answers from a photo and a voice recording you uploaded — proven by a question only those two files can answer.

There is no converter to set up: every project turns images and audio into text on the way in, with managed conversion.

  1. See the conversion rules your project already holds.
  2. Upload the photo.
  3. Upload the recording.
  4. Ingest the photo.
  5. Ingest the recording.
  6. Wait until both are ready.
  7. Read what conversion wrote.
  8. Point the agent at the media.
  9. Ask what only the media knows.

Every step is one API call, shown for all three clients. The ids in the responses are examples — copy the ones your own calls return.

Prerequisites​

  1. An agent with retrieval. Answer from your documents builds one and shows how retrieval works. Arrive here with:

    export NATURALI_TOKEN=nat_sk_...
    export PROJECT=proj_cT9LACJi0WypPf5U
    export AGENT=agent_lD7mNEur9S1cIuAQ
  2. Two files of your own in the current directory:

    • receipt.png — a photo or screenshot of any receipt;
    • meeting.mp3 — a short voice memo with one fact in it.

    The responses below come from a café receipt and a one-sentence recording, "Launch is next Tuesday." Yours will read differently; ask about what your files say.

Each conversion is a generation on a model naturali provides, billed from your credit balance, and a project with a negative balance cannot ingest. The converted documents count towards your account's storage allowance.

1. See the conversion rules​

Conversion is three ingestion rules naturali keeps in every project — one each for audio/*, image/* and application/pdf (scans only). List them with GET /v1/projects/{project_id}/ingestion-rules:

naturali list-ingestion-rules --project-id "$PROJECT"
{
"data": [
{
"id": "igr_ASGQESjiW8Bl5Fr2",
"content_type_glob": "audio/*",
"agent_id": "agent_WXn4tGuFzzgnGKJj",
"chunk_strategy": "whole",
"metadata": { "naturali_managed": "conversion", "template_version": 1 }
},
{
"id": "igr_1NeKM3KrWrHY7XPj",
"content_type_glob": "image/*",
"agent_id": "agent_WXn4tGuFzzgnGKJj",
"chunk_strategy": "whole",
"metadata": { "naturali_managed": "conversion", "template_version": 1 }
},
{
"id": "igr_gceI7aE0gDsSVhml",
"content_type_glob": "application/pdf",
"agent_id": "agent_WXn4tGuFzzgnGKJj",
"chunk_strategy": null,
"metadata": { "naturali_managed": "conversion", "template_version": 1 }
}
]
}

metadata.naturali_managed marks them: they are read-only, and a rule of your own on a narrower type, such as image/png, wins over them. No naturali_managed rules means conversion is off for the project — see Managed conversion to turn it back on.

2. Upload the photo​

The content_type is what picks the rule, so send the real one.

naturali upload-file-base64 \
--project-id "$PROJECT" \
--content "$(base64 -w0 receipt.png)" \
--filename receipt.png \
--prefix /media \
--content-type image/png
{
"id": "file_oKGatOttNow3AAw4",
"prefix": "/media",
"filename": "receipt.png",
"path": "/media/receipt.png",
"content_type": "image/png",
"size": 5863
}
export RECEIPT=file_oKGatOttNow3AAw4

3. Upload the recording​

naturali upload-file-base64 \
--project-id "$PROJECT" \
--content "$(base64 -w0 meeting.mp3)" \
--filename meeting.mp3 \
--prefix /media \
--content-type audio/mpeg
{
"id": "file_6ICGzQzyyOtCKIeD",
"prefix": "/media",
"filename": "meeting.mp3",
"path": "/media/meeting.mp3",
"content_type": "audio/mpeg",
"size": 26880
}
export RECORDING=file_6ICGzQzyyOtCKIeD

4. Ingest the photo​

POST /v1/projects/{project_id}/documents/ingest is the same call as for a text file. The image/* rule sends the photo to the converter, and the text it reads back becomes the document.

naturali ingest-document \
--project-id "$PROJECT" \
--file-id "$RECEIPT" \
--path-prefix /media/
{
"id": "doc_O46t8imSw0p6b4kz",
"file_id": "file_oKGatOttNow3AAw4",
"path": "/media/receipt.png",
"content_type": "image/png",
"status": "pending",
"version": 1
}
export RECEIPT_DOC=doc_O46t8imSw0p6b4kz

5. Ingest the recording​

The audio/* rule transcribes it.

naturali ingest-document \
--project-id "$PROJECT" \
--file-id "$RECORDING" \
--path-prefix /media/
{
"id": "doc_Vigfer3NrGJ93rRD",
"file_id": "file_6ICGzQzyyOtCKIeD",
"path": "/media/meeting.mp3",
"content_type": "audio/mpeg",
"status": "pending",
"version": 1
}
export RECORDING_DOC=doc_Vigfer3NrGJ93rRD

6. Wait until both are ready​

Conversion runs in the background, so poll GET /v1/projects/{project_id}/documents/{document_id}/status for each document until it reads ready. Short files like these took a few seconds. Shown for the photo; run it again with $RECORDING_DOC.

naturali get-document-status \
--project-id "$PROJECT" \
--document-id "$RECEIPT_DOC"
{
"id": "doc_O46t8imSw0p6b4kz",
"status": "ready",
"chunk_count": 1,
"total_chunks": 1,
"total_pages": 1,
"progress": 100
}

An image or a recording becomes a single chunk. failed carries the reason in error.

7. Read what conversion wrote​

GET /v1/projects/{project_id}/documents/{document_id} returns the converted text in content — exactly what retrieval will hand the agent. Shown for the photo; run it again with $RECORDING_DOC.

naturali get-document \
--project-id "$PROJECT" \
--document-id "$RECEIPT_DOC"
{
"id": "doc_O46t8imSw0p6b4kz",
"path": "/media/receipt.png",
"content_type": "image/png",
"status": "ready",
"content": "Corner Cafe\nCoffee 3.50\nSandwich 8.00\n\nTotal amount: 11.50"
}

The recording's document read "content": "Launch is next Tuesday." If the text is wrong here, the agent's answer will be too — fix the source file, then re-run it with POST /v1/projects/{project_id}/documents/{document_id}/ingest.

8. Point the agent at the media​

Scope retrieval to /media/ with PATCH /v1/projects/{project_id}/agents/{agent_id}. List /handbook/ too if the agent should keep answering from the handbook.

naturali patch-agent \
--project-id "$PROJECT" \
--agent-id "$AGENT" \
--instructions 'Answer from the provided context only. When it does not cover the question, answer only: I do not know.' \
--knowledge-config '{ "document_paths": ["/media/"], "limit": 3 }'
{
"id": "agent_lD7mNEur9S1cIuAQ",
"knowledge_config": { "limit": 3, "document_paths": ["/media/"] },
"version": 3
}

The instruction says to answer only "I do not know": a looser one let the model append it to an answer it had already given.

9. Ask what only the media knows​

One question, two facts: the total is only in the photo, the launch date only in the recording.

naturali create-agent-generation \
--project-id "$PROJECT" \
--agent-id "$AGENT" \
--wait true \
--messages '[{"role":"user","content":"What did I spend at the cafe, and when is the launch?"}]'
{
"id": "gen_p8YcysDsFiqABrCx",
"trace_id": "trace_JdjUOnI5nYYAkas2",
"status": "completed",
"output": {
"model": "glm-4.7-flash",
"content": "You spent $11.50 at the cafe. The launch is next Tuesday.",
"finish_reason": "stop"
}
}

That is the value: $11.50 was read off a photo and "next Tuesday" was heard in a recording, and the agent answered from both with no text you typed.

What's next​