The data store is a standalone, autonomous content-graph and retrieval service. Point it at a corpus directory and it runs itself: a filesystem connector detects source files, acquires them with genuine provenance, parses them, gates and activates a single active parse per source, builds semantic annotations, and serves retrieval evidence from projections. Operators mostly observe. The admin routes are rare overrides for the cases where the autonomous pipeline needs a deliberate hand.
This document is the operator entry point: how the fabric works and how to run it. For the exact contracts and comprehensive detail, see:
- INSTALL.md — installation, models, and one-time storage setup.
- PROTOCOL.md — the HTTP contract of record (request/response shapes, status codes, error envelopes).
- QUICKSTART.md — condensed install/operation command reference.
- SPEC-SERVER.md, SPEC-CLIENT.md, ARCHITECTURE.md — the comprehensive design references.
The service is first run at commissioning. The behavior described here is stated from the code as built; it is not a warranty of runtime-verified behavior.
The pipeline is autonomous. Nothing in routine operation requires an operator call — you supply a corpus and the scheduler drives the rest on an adaptive cadence:
corpus files
│
▼
┌──────────┐ filesystem connector: full scan, never follows symlinks,
│ DETECT │ (mtime, size) prescreen against known state, atomic staging
└────┬─────┘
▼
┌──────────┐ importer writes canonical AcquisitionRecords + SourceLocations;
│ ACQUIRE │ raw bytes land in the content-addressed artifact store first
└────┬─────┘
▼
┌──────────┐ EPUB: in-process EPUB worker; plain text: text worker.
│ PARSE │ Imported bundles become canonical units; invalid bundles fail.
└────┬─────┘
▼
┌──────────────┐ fine chunks, lexical index, fine and context dense vectors,
│ PROJECTIONS │ persisted ColBERT windows, and derived views
└────┬─────────┘
▼
┌──────────────┐ one active parse per source; dominance checks can hold
│ GATE/ACTIVATE│ a valid candidate for disposition instead of activating it
└────┬─────────┘
├──────────► RETRIEVE: POST /query serves the active parse
▼
┌──────────────┐ a separate worker runs entity, relation, and summary chains;
│ ANNOTATE │ completed outputs commit durably
└────┬─────────┘
└──────────► projection worker publishes graph, summary, dense + ColBERT
Only one parse is ever active per source: unit reads honor the active-parse gate, so a non-active parse's units are never served.
Retrieval combines source-dense, lexical, and grouped graph/semantic annotation rankings. ColBERT evaluates source and matched annotation representations; final reranking receives canonical passages plus separately labeled matched context. Publication and retrieval do not wait for all annotation types to finish.
Embedding and annotation never operate on single canonical units. Four grains
stack, each a run of consecutive members of the grain below it in reading
order, packed greedily under a ColBERT-token cap from [indexing]
(config.example.toml is the reference for every cap):
- Fine chunks (
indexing.fine_max_tokens) pack evidence members; they feed the fine dense vectors and the lexical index and are stored inchunk_projections. - ColBERT windows (
indexing.colbert_max_tokens) pack fine chunks; their token matrices are stored incolbert_windowsand scored by MaxSim. - Context windows (
indexing.context_max_tokens) pack fine chunks; they feed the context dense vectors, the annotation source windows, and the annotation cohorts. - Excerpts (
indexing.excerpt_windowsconsecutive context windows) are the unit of one annotation call and exist only in provenance.
Evidence members are derived once, at the fine chunker: a text_block with
role heading is not evidence and contributes only to the section path; all
cells of one table_row form one member whose text is the cell texts joined by
a tab; every other evidence-bearing unit is one member. Reading order of fine
chunks is chunk_projections.chunk_index, the only order authority for every
higher grain. Runs cross section boundaries; the section path of a run's first
member is carried as metadata.
Every grain has canonical text (member texts joined by one blank line) and
model input (the section path joined by " / ", a blank line, then the canonical
text). Canonical text feeds hashes, provenance ranges, the annotation source
matcher, and citations; model input feeds embedding and annotation calls, and
the prefix is never part of any range. A run below indexing.min_fill_ratio
of its cap is filled by splitting the next member (fine grain only, at a
sentence, then word boundary) or, at document end, merged into its predecessor
when the combined run fits; the one permitted sub-minimum run is a final run
that cannot merge. Every persisted run records fragments: the Unicode scalar
ranges of each unit's evidence text it is sliced from.
Each annotation call consumes one excerpt: a run of consecutive context windows of the active parse. The model receives the excerpt's canonical text as its only source text, with the section path as a separate context line. Entity, relation, and summary work use separate chains. Within a chain, model requests run sequentially; the server validates each response before passing its output to the next request.
- Entities: one request returns
entitiesentries withname/entityTypetaken directly from the excerpt, names as written. Each name must ground in the excerpt under the shared fuzzy matcher; a name that does not is dropped and counted, never a chain failure. Exactname/entityTyperepeats collapse to one entry; the same name under two types keeps both. Empty required text and an acceptedNOT_AN_ENTITYtype fail validation. - Relations: select source statements, form relationship triples for each selected statement, then request supporting quotations for bounded batches of those relationships. Later calls include the original excerpt. Statements and quotations use a shared fuzzy matcher: Unicode lowercasing and alphanumeric filtering ignore case, punctuation, whitespace, and word boundaries; separate limits allow text repairs and source omissions. The model's cleaned text is retained. Matching does not establish semantic correctness. Nonempty-text validation and complete receipt-index accounting remain; empty quotation lists are permitted.
- Summaries: request one summary per excerpt and validate its response shape.
Each produced item is attributed to the excerpt fragments whose text supports it: an entity keeps the fragments matching its name, a relation keeps the fragments matching any of its evidence quotes, and a summary keeps every fragment. An item that matches no fragment keeps every fragment.
An empty entities array, or one whose every name fails grounding, is a
successful empty result: the worker records fresh coverage so the excerpt is not
retried merely because it has no annotations. Model judgments are not
independently verified for semantic accuracy.
Intermediate results stay in memory. A completed chain commits its annotation set and any memo entry together; a failed chain restarts without intermediate checkpoints, under the retry policy below. Prompt/schema changes alter memo identity for new or unfinished work, while completed coverage remains fresh. Existing annotations are not automatically reclassified or removed.
Start the service. The server binary is data-store-service. After inference
initialization, it serves HTTP during the configured
server.startup_delay_seconds window before initializing corpus storage or
starting workers; 0 disables the wait. Send data-store --config config.toml --rebuild-all during this window to
clear the old corpus before ordinary startup can modify it. Health, Operation
polling, rebuild-all, and shutdown remain available; other storage-dependent
requests return 503 and readiness stays false. Rebuild-all ends the countdown
immediately; after successful clearing, normal startup continues without waiting
out the remaining delay. Failed rebuilds keep storage paused. Startup logging
and admin-token publication remain active.
Set up storage. Run the server with --setup-storage to build the
fabric hot plane's SQLite schema. This is one deliberate operator
action. The runtime NEVER creates or migrates schema — all schema arrives through
this explicit setup step. Re-running it validates an existing compatible
database and leaves it untouched; an incompatible one (schema version or table
contract mismatch) is deleted with the whole {index_root}/fabric directory and
recreated, after logging the reason and the row counts being lost. See
INSTALL.md for the full procedure and the model artifacts required before
first start.
data-store-service --config config.toml --setup-storage
# → fabric storage schema ready at <index_root>/fabric/...If the fabric plane is missing or invalid at startup, the service still serves,
but reports ready=false and health explains why (run --setup-storage).
Admin token handoff. On every start the service generates an admin bearer
token, prints it in the startup handoff, and publishes it to an owner-only token
file (admin.token_file_path, required; .data-store-admin-token as shipped in
config.example.toml). The bundled
CLI reads this file to authenticate protected calls; a raw curl needs the token
as Authorization: Bearer <token>. The token file is one of several owner-only
secret files; INSTALL.md lists the full set (annotator and HTTP-backend
API-key files).
Readiness. The top-level ready flag is the conjunction of exactly two
readiness-critical components: inference and sync. Everything else reported by
health is diagnostic-only and never makes a running service report unavailable.
Annotation work. The worker runs the annotation request chains
over excerpts. Each call enables thinking, requests structured JSON without
streaming, and allows the completion tokens set by
models.annotator.max_completion_tokens, including reasoning.
Projection work. A separate synchronous worker publishes graph and summary inputs independently, and builds dense/ColBERT representations for individual annotations, combined annotations per context window, and the context windows themselves as source representations. It discovers existing committed annotations without regenerating them. Embedding failures retain annotation results and the previous valid publication.
Operator policies. The entity-match document controls graph-entry fuzzy
matching. Inspect vocabulary with --vocabulary <entity|relation> before editing
it. Both policy documents remain required, validated, content-hashed, and
versioned at startup. The annotator-naming document is not applied to the current
single-goal prompts; editing it does not change producer memo identity.
Annotation dry run. data-store-service --annotation-dry-run <N> parses the
corpus and is meant to sample the first <N> excerpts per source per type
(entity and relation; no summaries or embeddings). It currently cannot sample:
excerpts are runs of context windows, the dry-run pass stops before the
projection build that constructs them, and the planner returns an empty plan
(annotator_plan.windows_unpublished in the service log). Inspect whatever
annotations exist with --vocabulary <entity|relation> all: sampled parses
are not active. Normal startup adopts the parses without re-converting them
and completes ingestion.
PROTOCOL.md is the contract of record. This is the map.
Public (no bearer token):
| Method & path | Purpose |
|---|---|
POST /query |
Synchronous retrieval; returns ranked passages and their canonical EvidencePack. |
GET /v1/health |
Readiness and per-component diagnostics. |
GET /v1/monitor |
Current ingestion, annotation, publication, and model-call observations. |
GET /units/{unitId} |
One canonical unit (served only if its parse is active). |
GET /units/{unitId}/relationships |
Unit relationships (direction/type filters). |
GET /sources/{sourceId} |
Source locations and freshness. |
GET /sync/status |
Published sync/scheduler health. |
Protected (bearer token required):
| Method & path | Purpose |
|---|---|
POST /sources |
Register/ingest a source (async Operation). |
POST /sources/{sourceId}/parses |
Force a parse run (async Operation). |
POST /sources/{sourceId}/parses/{parseId}/activate |
Activate a parse (async Operation). |
POST /parses/{parseId}/accept |
Accept a held parse (async Operation). |
POST /parses/{parseId}/discard |
Discard a held parse (async Operation). |
POST /snapshots |
Create a forensic snapshot (async Operation). |
POST /restore |
Restore a deactivated source's parse from its pre_deactivation snapshot and reactivate the source (async Operation). |
POST /rebuild-all |
Clear indexed corpus state and schedule automatic rebuilding (async Operation). |
POST /clear-failures |
Unblock failed background work while preserving successful work and failure history (async Operation). |
POST /shutdown |
Graceful shutdown (immediate confirmation, then signal). |
GET /parses?status=held |
List parses awaiting disposition. |
GET /operations/{operationId} |
Poll an async Operation's status. |
GET /annotations/vocabulary?annotationType=…&scope=… |
Inspect the grouped entity/relation annotation vocabulary. |
The service's client-facing API uses JSON responses and Operation polling. Internal annotator model calls return complete responses without streaming.
Every mutating admin route runs its work asynchronously. The route returns
202 Accepted with an operation id:
{ "operationId": "op_..." }Poll GET /operations/{operationId} until the Operation reaches a terminal
status (succeeded or failed). POST /shutdown is the one exception: it is a
control action, not async work, so it confirms immediately and is not an
Operation row.
Operation-succeeded is not the same as a parse outcome. An Operation reaching
succeeded means the pipeline lifecycle completed — the work ran to its end
without an internal error. It does NOT tell you the domain verdict. Whether a
parse was activated, held, or recorded a failed parse lives in the
parse run, not the Operation. To learn the verdict, consult the parse run
directly or list held candidates with GET /parses?status=held.
GET /v1/health returns:
{
"service": "data-store",
"ready": true,
"components": [
{
"name": "sync",
"ready": true,
"details": ["role: readiness-critical", "..."],
"counts": []
},
{
"name": "fabric",
"ready": true,
"details": ["role: diagnostic-only", "..."],
"counts": [
{ "label": "held", "source_system": "filesystem", "value": 0, "as_of": "2026-..." }
]
}
]
}Each component carries typed details and a typed counts array. Every count
carries its own as_of label — a count is never presented as current without
saying when it was measured — and fabric counts are keyed by source_system.
/v1/health and /v1/monitor use snake_case response fields.
Components that publish no counters serialize an empty array.
- Readiness-gating components:
inference,sync. Their combined readiness is the top-levelready. - Diagnostic-only components:
logging, and the fabric diagnosticsfabric,annotation,projections, andsearch_admission. These are visible for operators but never gate readiness. Thefabricandannotationcounters (held, serving-stale, stuck-building, access-lost, unparseable-mime, verification-halted, annotation freshness, retry exhaustion, and so on) are surfaced for observation only.
--health presents a compact operational report, with attention items and
unreported measurements first. Corpus exceptions appear once per source system;
zero-valued categories collapse into one line. --health-details retains every
component detail, startup smoke result, and diagnostic counter. Both commands
read the same endpoint; summaries come from the server's component snapshots.
Both health commands show each document in the worker's measured inventory:
completed / total (percentage), entity/relation/summary breakdowns, pending,
running, failed, retry-waiting, and exhausted work, and activity including commit
and storage waits. Source/parse/plan identity and measurement times identify the
captured work. Unknown totals and no required work are explicit. Completion
counts committed fresh coverage, including empty results and memo reuse; 100%
does not assert retrieval projection publication. The worker updates progress
during processing; newly active or changed sources appear on subsequent discovery.
Projection health separately reports published, pending, and failed graph, summary, and embedding cohorts per source/parse, with activity and measurement times. CLI and web health displays retain this distinction from annotation completion.
Last-cycle counts remain historical: eligible missing work excludes retry-waiting
and exhausted items, so zero does not establish completion. Parked workers retain
their reason. --sync-status reports ingestion. Model initialization does not
establish live endpoint health.
For current model activity, read the file configured by logging.file_path
(logs/data-store.log as shipped): stage starts, completions, measured
usage, failures, retry delays, and exhaustion are recorded there. Missing token
usage remains unknown; character counts are not token estimates. Existing
annotation log entries include annotation_progress="completed / total (percentage)".
Annotator requests omit response_format and stream from the transcript while
retaining temperature. Responses show only content, reasoning, completion tokens,
reasoning tokens, prompt tokens, and total tokens; missing fields are unavailable.
Content and reasoning are retained in full. The transcript is written to
logs/annotator.log, relative to the config directory. It appends
one readable group per call as calls finish, for both normal work and dry runs,
independently of the service-log level. Match call_id across both logs. Each
group contains REQUEST, RESPONSE, and RESULT together, including structural
validation or the failure/cancellation reason. Unfinished groups are buffered in
memory and can be lost on process termination; call starts remain in the service
log. Calls use stream: false; no chunk or generation-
progress records are written. A timeout or cancellation before body receipt is
reported without a partial response. Database commits remain
separate service-log events. Transcript open/write failures appear in the service
log; authentication credentials are excluded from the transcript.
The transcript shows document progress immediately before END CALL. Its
pre-persistence timing is preserved: calls in a wave may repeat the count, and
the final transcript entry may remain below 100%. Health and service-log commit
entries show subsequent completion. Unmeasured progress, including dry runs,
is unavailable; neither log adds entries for progress.
These are recorded MVP narrowings, not defects:
POST /queryomitsqueryExecutionRecordId. The QueryExecutionRecord audit tier is deferred post-MVP, so no QER is written and the response has noqueryExecutionRecordId. Any per-query id on the pack is a correlation handle, not a QER id.- The
multi_vectorretrieval channel is deferred post-MVP. The active retrieval channels are lexical, dense, graph, and semantic; graph and semantic matches share one outer fusion contribution. (ColBERT MaxSim is used internally as a reranking stage, not as a retrieval channel.) - Replay claims at MVP: evidence replay is
bit_exact; retrieval and generation replay arenot_supported. The system does not claim a replay fidelity it cannot demonstrate. - The query envelope is narrowed. The request DTO rejects unknown fields, so
the deferred spec query fields (
contentTypes,timeRange,metadataFilters,sourceSystems,freshness,channels,rerank,includeContradictions,includeFreshnessMetadata) are rejected if named. They can be added back additively.
The bundled data-store CLI wraps the HTTP surface, reads the admin token file
for protected calls, and renders results for operators. Its verbs map directly to
the routes above. Run data-store --help for the full list; the interactive REPL
is documented in SPEC-CLIENT.md §1.2.
Operator verbs (CLI flag / REPL name):
| Verb | Arguments | What it does |
|---|---|---|
--health |
— | Show compact operational health, prioritizing attention items. |
--health-details |
— | Show every component detail and diagnostic counter. |
--query |
<queryText...> |
Run a query; remaining args join into the query text. |
--query-raw |
<queryText...> |
Run the same query and print the complete response JSON. |
--ingest |
<sourceSystem> <nativeUri> |
Register/ingest a source. |
--reparse |
<sourceId> <sourceSystem> <nativeUri> |
Force a parse run. |
--activate |
<sourceId> <parseId> |
Activate a parse. |
--accept |
<parseId> |
Accept a held parse. |
--discard |
<parseId> |
Discard a held parse. |
--snapshot |
[requestJson] |
Create a snapshot. |
--restore |
<sourceId> <parseId> |
Restore a deactivated source's parse from its pre_deactivation snapshot and reactivate it. |
--rebuild-all |
— | Clear indexed state and artifacts, then automatically reingest the corpus. Available during the startup delay. |
--clear-failures |
— | Requeue failed work and reset annotation retry budgets without rebuilding successful work. |
--held-parses |
— | List held parses awaiting disposition. |
--operation |
<operationId> |
Read an Operation once. |
--vocabulary (--vocab) |
<entity|relation> [active|all] |
Inspect the annotation vocabulary (scope defaults to active). |
--unit |
<unitId> |
Read one unit. |
--relationships |
<unitId> [direction] [relationshipType] |
Read unit relationships. |
--source |
<sourceId> |
Read source locations and freshness. |
--sync-status |
— | Read sync/scheduler health. |
--shutdown |
— | Request graceful shutdown. |
For admin verbs the CLI submits the request, then polls the returned Operation until it terminates, and points you at the parse run / held-parses for the domain verdict.
--monitor and --serve are startup modes of the same binary, not verbs: see
"Ingestion monitor" and "Web UI" below.
queryText is the only required field. retrieval.default_results and
retrieval.max_results configure the default and maximum passage counts.
maxFinalEvidenceUnits selects a count within that range. Raw unit
locators default on; relationships, annotations, and debug default off.
curl -s http://127.0.0.1:8091/query \
-H 'Content-Type: application/json' \
-d '{
"queryText": "how does activation gating work",
"constraints": { "governanceDomains": ["local-corpus"] },
"retrievalPolicy": { "maxFinalEvidenceUnits": 5 },
"evidencePolicy": { "includeRelationships": true, "includeAnnotations": true },
"debug": false
}'The response carries ranked results with passage text, source locations,
and section headings, alongside the complete
canonical constituents in evidencePack. Oversized single-unit excerpts are
labeled truncated; their full bodies remain in the pack. Per-stage retrieval
diagnostics are attached only when debug is true.
Via the CLI, pass bare query text — the client builds the {"queryText": ...}
body itself, so the remaining arguments are the text (quoted or unquoted):
data-store --config config.toml --query how does activation gating workThe CLI and web console show each passage's matching search channels and
annotation contribution, including annotation representations, source ranges,
matched entities, and directed relationships.
This attribution is available without debug; it describes candidate matches,
not a measured improvement in retrieval quality. Web retrieval details retain
unit mappings and supporting references.
Dense retrieval combines fine-chunk and context-window matches. Individual and combined annotations are searched separately and share a grouped annotation ranking with graph matches. ColBERT MaxSim scores ColBERT windows; graph and annotation hits reach windows through shared fine-chunk membership, and passages are built from window fragments, so exact source excerpts remain identifiable through scoring, passage construction, reranking, and citations.
Query scoring streams persisted vectors with bounded buffers; operating-system
file caching can use spare RAM without requiring resident vector planes. Chunk
mappings and canonical metadata still scale with the captured corpus. Existing
annotations are embedded during normal discovery. Parses lacking required section
embeddings still need data-store --config config.toml --rebuild-all from the
project root, which clears indexed state and reingests the corpus. Snapshots
lacking the required passage/section representations are rejected before restore.
[indexing] sets the fine, ColBERT, and context caps, the excerpt length, and
the minimum fill (see "Retrieval grains"). Displayed passages use
retrieval.passage_max_tokens. Final reranking uses the candidate depth
retrieval.reranker_candidate_pool_size, raised for larger valid result
requests. Matched graph context shares the reranker's total input capacity
without a separate token cutoff.
The CLI prints each passage once with its citation. Use --query-raw (REPL:
query-raw) with the same query text to print the complete response JSON. Raw
output does not automatically enable request diagnostics.
curl -s http://127.0.0.1:8091/v1/healthdata-store --config config.toml --healthdata-store --config config.toml --monitorThe read-only terminal dashboard refreshes every 200 ms on one screen: ingestion
stages and queues, committed annotation coverage, retrieval publication, grouped
model calls, timings, reported token usage, waits, failures, and recent outcomes.
Progress bars use measured totals; files and deduplicated active known sources
remain separate. Blue, amber, and magenta accompany text state labels. Stale
connections and display overflow are explicit; resize for more space. Q, Esc,
or Ctrl-C exits. There are no monitor configuration settings or submenus.
See SPEC-CLIENT.md §1.7 for polling and terminal behavior.
The headline distinguishes service readiness from ingestion completion and explains blocked work. Annotation and publication percentages cover active documents only. Persisted parse and queue failures remain visible after restart; routine checks leave idle panels unchanged.
To unblock failed work after addressing its cause:
data-store --config config.toml --clear-failuresThis preserves successful work and failure history, resets retry eligibility, and resumes background processing. Command success does not mean ingestion is complete. Held parses and validation checks remain in force; workers terminated by panic or failed startup still require a restart.
--serve <host>:<port> is a startup mode of the same binary, not a verb: it runs
a local HTTP server in the foreground until the process is terminated, printing
the bound URL. --config still applies — serve mode reuses the client's service
base URL and admin token file.
data-store --config config.toml --serve 127.0.0.1:8092The UI offers a query console (constraints, passage limit, evidence toggles, debug) showing cited passages with raw evidence and diagnostics in details panels, a unit explorer with relationship filters, a source view, a health dashboard, sync status, the held-parses list, an operation viewer, and the vocabulary explorer.
v1 is read-only. The browser reaches the service only through a fixed allowlist
of proxy routes under /api/*, which exposes no mutating route; admin reads use
the client's admin token file, and the UI itself carries no authentication.
Configuration is a single TOML file (config.toml, from config.example.toml).
Missing required settings and unknown keys are fatal startup errors. Configuration
comments specify units, enforcement scope, and excess-input behavior. The sections:
| Section | Purpose |
|---|---|
[server] |
HTTP bind address, required nonnegative integer startup_delay_seconds (0 disables the wait), and request-shape limits. |
[logging] |
File-backed service logging. |
[admin] |
Admin token file location. |
[client] |
Client deadlines, polling cadence, and provenance previews. |
[retrieval], [retrieval.entity_matching] |
Result counts, candidate depths, graph matching, passage/evidence budgets, and query admission. |
[indexing] |
ColBERT-token caps for fine chunks, ColBERT windows, and context windows; context windows per excerpt; minimum fill ratio. |
[resources] |
Read, allocation, and inventory admission guards. |
[workers] |
Annotation/projection batching, concurrency, polling, and publication allowances. |
[sqlite] |
Lock-wait timeout, cooperative SQL execution timeout, and progress-callback cadence. |
[parsing] |
Candidate unit, relationship, and warning caps and the unit-body byte limit. |
[scheduling] |
Sync cadence, backoff, and maintenance polling. |
[diagnostics] |
Operational error, identifier, and progress-summary bounds. |
[inference] |
Accelerator selection for local retrieval models. Required but unused when dense, ColBERT, and the reranker all use HTTP; no local accelerator is initialized in that mode. |
[storage] |
Corpus root and service-owned index root. |
[connectors.filesystem] |
Governance domain stamped on acquired sources. |
[epub] |
EPUB admission budgets, all required: archive member count, per-member and total decompressed bytes, XML document bytes, image bytes, and element nesting depth. |
[models] |
Dense, ColBERT, and reranker each select an exclusive local or http backend. Remote ColBERT uses vLLM /pooling token inference, a matching local tokenizer, persisted document matrices, and CPU MaxSim; no local ColBERT weights are loaded. The annotator uses an external chat-completions endpoint. See INSTALL.md for backend fields. |
[policies] |
Paths to entity-match enable flags and annotator naming rules. Numeric entity-matching limits live in config.toml. |
See INSTALL.md for the annotated example and the required absolute paths.
Model serving capacity is separate from grain size. models.dense.max_tokens
and models.reranker.max_tokens declare the served capacities, and
models.colbert.document_max_tokens the ColBERT document limit; startup
requires /v1/models to advertise the configured capacity, requires
indexing.colbert_max_tokens to equal the ColBERT document limit, and requires
a context window plus its section-path prefix budget to fit the dense and
reranker capacities. Server-side truncation is disabled; oversized HTTP inputs
fail visibly. A ColBERT window whose prefixed input would exceed the document
limit is embedded without the prefix.
Construction settings are recorded with new projections. Restore validates their
recorded settings or explicit legacy formats without inference. Current resource
guards may refuse large historical artifacts without declaring them corrupt.
[parsing] and [epub] admission limits do not change parser identities.
The EPUB worker parses .epub files (EPUB 2 and EPUB 3) in-process with no
external tool. It records only structure the source declares: sections from
the navigation document or NCX and from headings, print pages from declared
page-break markers, lists, tables decomposed to cells, figures with their
image bytes archived by hash, captions, code blocks from pre, asides, and
footnote and cross-reference links. No structure is inferred from text
content; a source that declares no semantics for a feature gets no such
feature. DRM-protected archives are recorded parse failures. Other formats are
converted to plain text outside this service; only text/plain and EPUB are
routed.
These [models.annotator] settings are required; config.example.toml holds
the shipped values. Excerpt size is not an annotator setting: it follows from
indexing.context_max_tokens and indexing.excerpt_windows.
| Setting | Meaning |
|---|---|
timeout_seconds |
Deadline for each model call, including thinking. |
max_completion_tokens |
Completion-token allowance per call, including reasoning. |
annotation_max_retries |
Malformed-output retries after the initial attempt. |
annotation_retry_interval_seconds |
Fixed malformed-output retry interval, with no backoff ceiling. |
execution_max_retries |
Execution-failure retries after the initial attempt. |
execution_retry_initial_delay_seconds |
Initial execution-failure retry delay. |
execution_retry_max_delay_seconds |
Execution-failure backoff ceiling. |
The two failure counters are independent. Execution failures include timeouts,
HTTP/protocol errors, token-limit termination, and internal producer failures.
Their delay doubles up to the configured ceiling; it is independent of the
configured scheduler waits and dense HTTP retry backoff. Malformed outputs use their fixed interval.
Both limits allow 0 to disable that category's retries. Delays must be positive,
and the execution ceiling must be at least the initial delay.
Temperature starts at 0.0 and becomes
min(malformed_output_failures / annotation_max_retries, 1.0) on retries;
execution failures do not advance it. With a zero annotation retry allowance,
no temperature division or malformed-output retry occurs.
Once either counter exceeds its allowance, the annotation remains failed and is skipped as exhausted. Cancellation and scheduling deferrals spend neither budget. Counters and eligibility timers reset on restart or rebuild; failed chains restart as a whole, without intermediate-stage checkpoints.