GUIDES
The knowledge base
Give the agent something to read. Upload your docs once, and every answer is grounded in them.
Upload docs in the console
Open your project, go to Knowledge, and drop files in. Agentifys extracts the text, splits it into chunks, embeds them, and stores the vectors alongside the raw text. Indexing runs in the background. A doc goes live for retrieval the moment its job finishes, usually a few seconds after upload.
Console uploads default to the shared project pool, which answers every visitor to your widget. Product docs, policies, FAQs, pricing, release notes, put whatever you want the agent to know here. The upload form also has a Specific users visibility option, which writes the same document straight into one or more private per-user pools instead.
| Field | Type | Description |
|---|---|---|
Supported types | .pdf / .txt / .md | Checked by file extension. Anything else is rejected with UNSUPPORTED_TYPE. |
File size | <= 50 MB | Per file. Larger files are rejected with FILE_TOO_LARGE. |
Per-user cap | 20 docs | Applies to private per-user uploads through the API (below). Console uploads have no document cap. |
Languages | ar / en / mixed | Embeddings are multilingual, so an Arabic question can match English source text and back. |
Chunking is fixed, not configurable. Text is split per page at 400 words with an 80-word overlap, and any chunk shorter than 30 words is discarded as noise.
| Field | Type | Description |
|---|---|---|
Chunk size | 400 words | Word-level, never crossing a page boundary. PDFs keep their real page numbers for citation. |
Overlap | 80 words | Each chunk starts 320 words after the previous one, so a sentence split across a boundary still appears whole in one chunk. |
Minimum chunk | 30 words | Shorter chunks are dropped. A file made only of short lines can produce zero chunks and fail to index. |
.txt file. A file that yields text but no chunk of 30+ words fails with "Document is too sparse to index" and the total word count; merge short files into one before uploading.How retrieval works
Every question runs a vector search and a full-text search, but the two are not equal halves. The candidate pool is the vector top 8 and nothing else. Postgres full-text search runs over the same chunks and is then joined onto those 8 rows to adjust their order.
| Field | Type | Description |
|---|---|---|
Embedding | BAAI/bge-m3 | Multilingual, 1024-dimensional. The same model embeds the chunk at upload and the question at query time. |
Candidates | 8 chunks | The pgvector cosine top 8, scoped to the audiences the caller may see. This is the entire candidate pool. |
Keyword side | tsvector 'simple' | Language-agnostic Postgres FTS over the same chunks, left-joined onto the 8 candidates. It re-orders them; it cannot add one. |
Ordering | weighted sum | Candidates are ordered by vec_score * 0.7 + text_score * 10 * 0.3 before reranking. |
Rerank | cross-encoder | mmarco-mMiniLMv2-L12-H384-v1 re-scores all 8 against the question. The raw score passes through a sigmoid, so reranker_score lands in 0..1. |
Gates | 0.30 and 0.30 | A chunk survives only if vec_score >= 0.30 AND reranker_score >= 0.30. Both floors apply. |
Injected | top 4 chunks | Survivors are sorted by reranker score, deduplicated on their first 120 characters, and the best four go into context. |
The upshot: when your docs contain the answer in language close to the question, it lands in context. When they do not, nothing clears the two 0.30 floors and the agent answers from the model instead of inventing a citation.
rag_retrieve record holding the first 120 characters of the question, how many chunks were injected, the top reranker score, the retrieval latency, and the 0.30 threshold. To read the actual text and per-chunk scores, use Knowledge → search in the console: it returns every candidate with its vec_score, reranker_score and a passed flag, and it can run the search as a specific user so you can preview what their private pool adds.passed: true there and still never reach the model, because chat only ever looks at the top 8. Treat the console as a diagnostic: a chunk sitting low in that list is a chunk chat will probably not see.Private per-user uploads
Retrieval runs against audiences. The shared project pool answers everyone. A private user:<id> pool answers only that one person. That is how you let a customer upload their own contract or spreadsheet and have the agent reason over it, without leaking it to the next visitor.
The private audience is only added to a user's retrieval when the project is in signedidentity mode and the id is neither empty nor anonymous. In open mode the caller-supplied id is forgeable, so retrieval stays project-only and a private pool is never read. Uploads require signed identity for the same reason.
DATABASE_URL) an upload with a user audience returns an error, and a per-user list returns an empty array rather than leaking the project pool. Only the shared project pool works there.Private uploads go through the API, not the console, and they require signed identity so the file is bound to a real user you vouched for. Post the file and the signed user_context as multipart form data.
/v1/knowledge/v1/knowledge call is origin-checked, including server-to-server calls. A request with no Origin header fails with 403 ORIGIN_NOT_ALLOWED, so a plain backend curl that works against /v1/credentials will not work here. Send an Origin header matching one of your allowed origins (a trailing slash on the header is ignored before comparison), or leave the allowlist empty.# Origin is required whenever the project has an allowed-origins list.
curl -X POST https://agentifys.ai/v1/knowledge \
-H "Authorization: Bearer era_your_project_key" \
-H "Origin: https://app.example.com" \
-F "file=@contract.pdf" \
-F 'user_context={"id":"u_123","_ts":1735689600,"_sig":"<hmac>"}'The call returns immediately while indexing runs in the background. Job and document ids are bare 12-character hex strings with no prefix:
{
"job_id": "9f3c1e7a2b8d",
"status": "processing",
"filename": "contract.pdf"
}List or delete a user's own docs with the same signed identity:
/v1/knowledge/v1/knowledge/{doc_id}SIGNED_IDENTITY_REQUIRED; a missing or anonymous id gives USER_ID_REQUIRED. Each user gets 20 docs; the 21st returns DOC_CAP_REACHED, and an unsupported file returns UNSUPPORTED_TYPE.Uploads are throttled per user, on a bucket separate from /v1/chat: 10 uploads a minute and 50 a day per user per project. Over either limit the call returns 429 RATE_LIMIT_EXCEEDED with a retry hint. Listing and deleting are not throttled. Indexing also has a global kill switch; while it is off, uploads return 503 INDEXING_PAUSED and search over already-indexed documents keeps working.
The widget already has the upload UI
You do not have to build an upload screen. The embedded widget ships one. A paperclip button appears in the composer on its own once two things are true: the project is in signed identity mode with indexing enabled, and ERAConfig.user.id is set. It accepts .pdf, .txt and .md, posts to /v1/knowledge, polls the job, and then shows "filename is ready — ask me about it" with a Remove link that calls the delete endpoint.
Reach for the API directly when the document comes from your own systems rather than from the user: a contract you already hold, an export you generate on their behalf, a file they uploaded elsewhere in your product.