Skip to content

Latest commit

 

History

173 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Data Store — Canonical Content-Graph and Retrieval Fabric

The data store is a standalone, autonomous content-graph and retrieval service. Point it at a corpus directory and it runs itself: a filesystem connector detects source files, acquires them with genuine provenance, parses them, gates and activates a single active parse per source, builds semantic annotations, and serves retrieval evidence from projections. Operators mostly observe. The admin routes are rare overrides for the cases where the autonomous pipeline needs a deliberate hand.

This document is the operator entry point: how the fabric works and how to run it. For the exact contracts and comprehensive detail, see:

  • INSTALL.md — installation, models, and one-time storage setup.
  • PROTOCOL.md — the HTTP contract of record (request/response shapes, status codes, error envelopes).
  • QUICKSTART.md — condensed install/operation command reference.
  • SPEC-SERVER.md, SPEC-CLIENT.md, ARCHITECTURE.md — the comprehensive design references.

The service is first run at commissioning. The behavior described here is stated from the code as built; it is not a warranty of runtime-verified behavior.

How the fabric works

The pipeline is autonomous. Nothing in routine operation requires an operator call — you supply a corpus and the scheduler drives the rest on an adaptive cadence:

  corpus files
       │
       ▼
  ┌──────────┐   filesystem connector: full scan, never follows symlinks,
  │  DETECT  │   (mtime, size) prescreen against known state, atomic staging
  └────┬─────┘
       ▼
  ┌──────────┐   importer writes canonical AcquisitionRecords + SourceLocations;
  │ ACQUIRE  │   raw bytes land in the content-addressed artifact store first
  └────┬─────┘
       ▼
  ┌──────────┐   EPUB: in-process EPUB worker; plain text: text worker.
  │  PARSE   │   Imported bundles become canonical units; invalid bundles fail.
  └────┬─────┘
       ▼
  ┌──────────────┐   fine chunks, lexical index, fine and context dense vectors,
  │ PROJECTIONS  │   persisted ColBERT windows, and derived views
  └────┬─────────┘
       ▼
  ┌──────────────┐   one active parse per source; dominance checks can hold
  │ GATE/ACTIVATE│   a valid candidate for disposition instead of activating it
  └────┬─────────┘
       ├──────────► RETRIEVE: POST /query serves the active parse
       ▼
  ┌──────────────┐   a separate worker runs entity, relation, and summary chains;
  │  ANNOTATE    │   completed outputs commit durably
  └────┬─────────┘
       └──────────► projection worker publishes graph, summary, dense + ColBERT

Only one parse is ever active per source: unit reads honor the active-parse gate, so a non-active parse's units are never served.

Retrieval combines source-dense, lexical, and grouped graph/semantic annotation rankings. ColBERT evaluates source and matched annotation representations; final reranking receives canonical passages plus separately labeled matched context. Publication and retrieval do not wait for all annotation types to finish.

Retrieval grains

Embedding and annotation never operate on single canonical units. Four grains stack, each a run of consecutive members of the grain below it in reading order, packed greedily under a ColBERT-token cap from [indexing] (config.example.toml is the reference for every cap):

  • Fine chunks (indexing.fine_max_tokens) pack evidence members; they feed the fine dense vectors and the lexical index and are stored in chunk_projections.
  • ColBERT windows (indexing.colbert_max_tokens) pack fine chunks; their token matrices are stored in colbert_windows and scored by MaxSim.
  • Context windows (indexing.context_max_tokens) pack fine chunks; they feed the context dense vectors, the annotation source windows, and the annotation cohorts.
  • Excerpts (indexing.excerpt_windows consecutive context windows) are the unit of one annotation call and exist only in provenance.

Evidence members are derived once, at the fine chunker: a text_block with role heading is not evidence and contributes only to the section path; all cells of one table_row form one member whose text is the cell texts joined by a tab; every other evidence-bearing unit is one member. Reading order of fine chunks is chunk_projections.chunk_index, the only order authority for every higher grain. Runs cross section boundaries; the section path of a run's first member is carried as metadata.

Every grain has canonical text (member texts joined by one blank line) and model input (the section path joined by " / ", a blank line, then the canonical text). Canonical text feeds hashes, provenance ranges, the annotation source matcher, and citations; model input feeds embedding and annotation calls, and the prefix is never part of any range. A run below indexing.min_fill_ratio of its cap is filled by splitting the next member (fine grain only, at a sentence, then word boundary) or, at document end, merged into its predecessor when the combined run fits; the one permitted sub-minimum run is a final run that cannot merge. Every persisted run records fragments: the Unicode scalar ranges of each unit's evidence text it is sliced from.

Annotation request chains

Each annotation call consumes one excerpt: a run of consecutive context windows of the active parse. The model receives the excerpt's canonical text as its only source text, with the section path as a separate context line. Entity, relation, and summary work use separate chains. Within a chain, model requests run sequentially; the server validates each response before passing its output to the next request.

  • Entities: one request returns entities entries with name/entityType taken directly from the excerpt, names as written. Each name must ground in the excerpt under the shared fuzzy matcher; a name that does not is dropped and counted, never a chain failure. Exact name/entityType repeats collapse to one entry; the same name under two types keeps both. Empty required text and an accepted NOT_AN_ENTITY type fail validation.
  • Relations: select source statements, form relationship triples for each selected statement, then request supporting quotations for bounded batches of those relationships. Later calls include the original excerpt. Statements and quotations use a shared fuzzy matcher: Unicode lowercasing and alphanumeric filtering ignore case, punctuation, whitespace, and word boundaries; separate limits allow text repairs and source omissions. The model's cleaned text is retained. Matching does not establish semantic correctness. Nonempty-text validation and complete receipt-index accounting remain; empty quotation lists are permitted.
  • Summaries: request one summary per excerpt and validate its response shape.

Each produced item is attributed to the excerpt fragments whose text supports it: an entity keeps the fragments matching its name, a relation keeps the fragments matching any of its evidence quotes, and a summary keeps every fragment. An item that matches no fragment keeps every fragment.

An empty entities array, or one whose every name fails grounding, is a successful empty result: the worker records fresh coverage so the excerpt is not retried merely because it has no annotations. Model judgments are not independently verified for semantic accuracy.

Intermediate results stay in memory. A completed chain commits its annotation set and any memo entry together; a failed chain restarts without intermediate checkpoints, under the retry policy below. Prompt/schema changes alter memo identity for new or unfinished work, while completed coverage remains fresh. Existing annotations are not automatically reclassified or removed.

Operating model

Start the service. The server binary is data-store-service. After inference initialization, it serves HTTP during the configured server.startup_delay_seconds window before initializing corpus storage or starting workers; 0 disables the wait. Send data-store --config config.toml --rebuild-all during this window to clear the old corpus before ordinary startup can modify it. Health, Operation polling, rebuild-all, and shutdown remain available; other storage-dependent requests return 503 and readiness stays false. Rebuild-all ends the countdown immediately; after successful clearing, normal startup continues without waiting out the remaining delay. Failed rebuilds keep storage paused. Startup logging and admin-token publication remain active.

Set up storage. Run the server with --setup-storage to build the fabric hot plane's SQLite schema. This is one deliberate operator action. The runtime NEVER creates or migrates schema — all schema arrives through this explicit setup step. Re-running it validates an existing compatible database and leaves it untouched; an incompatible one (schema version or table contract mismatch) is deleted with the whole {index_root}/fabric directory and recreated, after logging the reason and the row counts being lost. See INSTALL.md for the full procedure and the model artifacts required before first start.

data-store-service --config config.toml --setup-storage
# → fabric storage schema ready at <index_root>/fabric/...

If the fabric plane is missing or invalid at startup, the service still serves, but reports ready=false and health explains why (run --setup-storage).

Admin token handoff. On every start the service generates an admin bearer token, prints it in the startup handoff, and publishes it to an owner-only token file (admin.token_file_path, required; .data-store-admin-token as shipped in config.example.toml). The bundled CLI reads this file to authenticate protected calls; a raw curl needs the token as Authorization: Bearer <token>. The token file is one of several owner-only secret files; INSTALL.md lists the full set (annotator and HTTP-backend API-key files).

Readiness. The top-level ready flag is the conjunction of exactly two readiness-critical components: inference and sync. Everything else reported by health is diagnostic-only and never makes a running service report unavailable.

Annotation work. The worker runs the annotation request chains over excerpts. Each call enables thinking, requests structured JSON without streaming, and allows the completion tokens set by models.annotator.max_completion_tokens, including reasoning.

Projection work. A separate synchronous worker publishes graph and summary inputs independently, and builds dense/ColBERT representations for individual annotations, combined annotations per context window, and the context windows themselves as source representations. It discovers existing committed annotations without regenerating them. Embedding failures retain annotation results and the previous valid publication.

Operator policies. The entity-match document controls graph-entry fuzzy matching. Inspect vocabulary with --vocabulary <entity|relation> before editing it. Both policy documents remain required, validated, content-hashed, and versioned at startup. The annotator-naming document is not applied to the current single-goal prompts; editing it does not change producer memo identity.

Annotation dry run. data-store-service --annotation-dry-run <N> parses the corpus and is meant to sample the first <N> excerpts per source per type (entity and relation; no summaries or embeddings). It currently cannot sample: excerpts are runs of context windows, the dry-run pass stops before the projection build that constructs them, and the planner returns an empty plan (annotator_plan.windows_unpublished in the service log). Inspect whatever annotations exist with --vocabulary <entity|relation> all: sampled parses are not active. Normal startup adopts the parses without re-converting them and completes ingestion.

The HTTP surface at a glance

PROTOCOL.md is the contract of record. This is the map.

Public (no bearer token):

Method & path Purpose
POST /query Synchronous retrieval; returns ranked passages and their canonical EvidencePack.
GET /v1/health Readiness and per-component diagnostics.
GET /v1/monitor Current ingestion, annotation, publication, and model-call observations.
GET /units/{unitId} One canonical unit (served only if its parse is active).
GET /units/{unitId}/relationships Unit relationships (direction/type filters).
GET /sources/{sourceId} Source locations and freshness.
GET /sync/status Published sync/scheduler health.

Protected (bearer token required):

Method & path Purpose
POST /sources Register/ingest a source (async Operation).
POST /sources/{sourceId}/parses Force a parse run (async Operation).
POST /sources/{sourceId}/parses/{parseId}/activate Activate a parse (async Operation).
POST /parses/{parseId}/accept Accept a held parse (async Operation).
POST /parses/{parseId}/discard Discard a held parse (async Operation).
POST /snapshots Create a forensic snapshot (async Operation).
POST /restore Restore a deactivated source's parse from its pre_deactivation snapshot and reactivate the source (async Operation).
POST /rebuild-all Clear indexed corpus state and schedule automatic rebuilding (async Operation).
POST /clear-failures Unblock failed background work while preserving successful work and failure history (async Operation).
POST /shutdown Graceful shutdown (immediate confirmation, then signal).
GET /parses?status=held List parses awaiting disposition.
GET /operations/{operationId} Poll an async Operation's status.
GET /annotations/vocabulary?annotationType=…&scope=… Inspect the grouped entity/relation annotation vocabulary.

The polling admin model

The service's client-facing API uses JSON responses and Operation polling. Internal annotator model calls return complete responses without streaming.

Every mutating admin route runs its work asynchronously. The route returns 202 Accepted with an operation id:

{ "operationId": "op_..." }

Poll GET /operations/{operationId} until the Operation reaches a terminal status (succeeded or failed). POST /shutdown is the one exception: it is a control action, not async work, so it confirms immediately and is not an Operation row.

Operation-succeeded is not the same as a parse outcome. An Operation reaching succeeded means the pipeline lifecycle completed — the work ran to its end without an internal error. It does NOT tell you the domain verdict. Whether a parse was activated, held, or recorded a failed parse lives in the parse run, not the Operation. To learn the verdict, consult the parse run directly or list held candidates with GET /parses?status=held.

Health

GET /v1/health returns:

{
  "service": "data-store",
  "ready": true,
  "components": [
    {
      "name": "sync",
      "ready": true,
      "details": ["role: readiness-critical", "..."],
      "counts": []
    },
    {
      "name": "fabric",
      "ready": true,
      "details": ["role: diagnostic-only", "..."],
      "counts": [
        { "label": "held", "source_system": "filesystem", "value": 0, "as_of": "2026-..." }
      ]
    }
  ]
}

Each component carries typed details and a typed counts array. Every count carries its own as_of label — a count is never presented as current without saying when it was measured — and fabric counts are keyed by source_system. /v1/health and /v1/monitor use snake_case response fields. Components that publish no counters serialize an empty array.

  • Readiness-gating components: inference, sync. Their combined readiness is the top-level ready.
  • Diagnostic-only components: logging, and the fabric diagnostics fabric, annotation, projections, and search_admission. These are visible for operators but never gate readiness. The fabric and annotation counters (held, serving-stale, stuck-building, access-lost, unparseable-mime, verification-halted, annotation freshness, retry exhaustion, and so on) are surfaced for observation only.

--health presents a compact operational report, with attention items and unreported measurements first. Corpus exceptions appear once per source system; zero-valued categories collapse into one line. --health-details retains every component detail, startup smoke result, and diagnostic counter. Both commands read the same endpoint; summaries come from the server's component snapshots.

Both health commands show each document in the worker's measured inventory: completed / total (percentage), entity/relation/summary breakdowns, pending, running, failed, retry-waiting, and exhausted work, and activity including commit and storage waits. Source/parse/plan identity and measurement times identify the captured work. Unknown totals and no required work are explicit. Completion counts committed fresh coverage, including empty results and memo reuse; 100% does not assert retrieval projection publication. The worker updates progress during processing; newly active or changed sources appear on subsequent discovery.

Projection health separately reports published, pending, and failed graph, summary, and embedding cohorts per source/parse, with activity and measurement times. CLI and web health displays retain this distinction from annotation completion.

Last-cycle counts remain historical: eligible missing work excludes retry-waiting and exhausted items, so zero does not establish completion. Parked workers retain their reason. --sync-status reports ingestion. Model initialization does not establish live endpoint health. For current model activity, read the file configured by logging.file_path (logs/data-store.log as shipped): stage starts, completions, measured usage, failures, retry delays, and exhaustion are recorded there. Missing token usage remains unknown; character counts are not token estimates. Existing annotation log entries include annotation_progress="completed / total (percentage)".

Annotator requests omit response_format and stream from the transcript while retaining temperature. Responses show only content, reasoning, completion tokens, reasoning tokens, prompt tokens, and total tokens; missing fields are unavailable. Content and reasoning are retained in full. The transcript is written to logs/annotator.log, relative to the config directory. It appends one readable group per call as calls finish, for both normal work and dry runs, independently of the service-log level. Match call_id across both logs. Each group contains REQUEST, RESPONSE, and RESULT together, including structural validation or the failure/cancellation reason. Unfinished groups are buffered in memory and can be lost on process termination; call starts remain in the service log. Calls use stream: false; no chunk or generation- progress records are written. A timeout or cancellation before body receipt is reported without a partial response. Database commits remain separate service-log events. Transcript open/write failures appear in the service log; authentication credentials are excluded from the transcript.

The transcript shows document progress immediately before END CALL. Its pre-persistence timing is preserved: calls in a wave may repeat the count, and the final transcript entry may remain below 100%. Health and service-log commit entries show subsequent completion. Unmeasured progress, including dry runs, is unavailable; neither log adds entries for progress.

Known deviations (stated where an operator meets them)

These are recorded MVP narrowings, not defects:

  • POST /query omits queryExecutionRecordId. The QueryExecutionRecord audit tier is deferred post-MVP, so no QER is written and the response has no queryExecutionRecordId. Any per-query id on the pack is a correlation handle, not a QER id.
  • The multi_vector retrieval channel is deferred post-MVP. The active retrieval channels are lexical, dense, graph, and semantic; graph and semantic matches share one outer fusion contribution. (ColBERT MaxSim is used internally as a reranking stage, not as a retrieval channel.)
  • Replay claims at MVP: evidence replay is bit_exact; retrieval and generation replay are not_supported. The system does not claim a replay fidelity it cannot demonstrate.
  • The query envelope is narrowed. The request DTO rejects unknown fields, so the deferred spec query fields (contentTypes, timeRange, metadataFilters, sourceSystems, freshness, channels, rerank, includeContradictions, includeFreshnessMetadata) are rejected if named. They can be added back additively.

Getting started

The bundled data-store CLI wraps the HTTP surface, reads the admin token file for protected calls, and renders results for operators. Its verbs map directly to the routes above. Run data-store --help for the full list; the interactive REPL is documented in SPEC-CLIENT.md §1.2.

Operator verbs (CLI flag / REPL name):

Verb Arguments What it does
--health — Show compact operational health, prioritizing attention items.
--health-details — Show every component detail and diagnostic counter.
--query <queryText...> Run a query; remaining args join into the query text.
--query-raw <queryText...> Run the same query and print the complete response JSON.
--ingest <sourceSystem> <nativeUri> Register/ingest a source.
--reparse <sourceId> <sourceSystem> <nativeUri> Force a parse run.
--activate <sourceId> <parseId> Activate a parse.
--accept <parseId> Accept a held parse.
--discard <parseId> Discard a held parse.
--snapshot [requestJson] Create a snapshot.
--restore <sourceId> <parseId> Restore a deactivated source's parse from its pre_deactivation snapshot and reactivate it.
--rebuild-all — Clear indexed state and artifacts, then automatically reingest the corpus. Available during the startup delay.
--clear-failures — Requeue failed work and reset annotation retry budgets without rebuilding successful work.
--held-parses — List held parses awaiting disposition.
--operation <operationId> Read an Operation once.
--vocabulary (--vocab) <entity|relation> [active|all] Inspect the annotation vocabulary (scope defaults to active).
--unit <unitId> Read one unit.
--relationships <unitId> [direction] [relationshipType] Read unit relationships.
--source <sourceId> Read source locations and freshness.
--sync-status — Read sync/scheduler health.
--shutdown — Request graceful shutdown.

For admin verbs the CLI submits the request, then polls the returned Operation until it terminates, and points you at the parse run / held-parses for the domain verdict.

--monitor and --serve are startup modes of the same binary, not verbs: see "Ingestion monitor" and "Web UI" below.

Example: query the fabric

queryText is the only required field. retrieval.default_results and retrieval.max_results configure the default and maximum passage counts. maxFinalEvidenceUnits selects a count within that range. Raw unit locators default on; relationships, annotations, and debug default off.

curl -s http://127.0.0.1:8091/query \
  -H 'Content-Type: application/json' \
  -d '{
        "queryText": "how does activation gating work",
        "constraints": { "governanceDomains": ["local-corpus"] },
        "retrievalPolicy": { "maxFinalEvidenceUnits": 5 },
        "evidencePolicy": { "includeRelationships": true, "includeAnnotations": true },
        "debug": false
      }'

The response carries ranked results with passage text, source locations, and section headings, alongside the complete canonical constituents in evidencePack. Oversized single-unit excerpts are labeled truncated; their full bodies remain in the pack. Per-stage retrieval diagnostics are attached only when debug is true.

Via the CLI, pass bare query text — the client builds the {"queryText": ...} body itself, so the remaining arguments are the text (quoted or unquoted):

data-store --config config.toml --query how does activation gating work

The CLI and web console show each passage's matching search channels and annotation contribution, including annotation representations, source ranges, matched entities, and directed relationships. This attribution is available without debug; it describes candidate matches, not a measured improvement in retrieval quality. Web retrieval details retain unit mappings and supporting references.

Dense retrieval combines fine-chunk and context-window matches. Individual and combined annotations are searched separately and share a grouped annotation ranking with graph matches. ColBERT MaxSim scores ColBERT windows; graph and annotation hits reach windows through shared fine-chunk membership, and passages are built from window fragments, so exact source excerpts remain identifiable through scoring, passage construction, reranking, and citations.

Query scoring streams persisted vectors with bounded buffers; operating-system file caching can use spare RAM without requiring resident vector planes. Chunk mappings and canonical metadata still scale with the captured corpus. Existing annotations are embedded during normal discovery. Parses lacking required section embeddings still need data-store --config config.toml --rebuild-all from the project root, which clears indexed state and reingests the corpus. Snapshots lacking the required passage/section representations are rejected before restore.

[indexing] sets the fine, ColBERT, and context caps, the excerpt length, and the minimum fill (see "Retrieval grains"). Displayed passages use retrieval.passage_max_tokens. Final reranking uses the candidate depth retrieval.reranker_candidate_pool_size, raised for larger valid result requests. Matched graph context shares the reranker's total input capacity without a separate token cutoff.

The CLI prints each passage once with its citation. Use --query-raw (REPL: query-raw) with the same query text to print the complete response JSON. Raw output does not automatically enable request diagnostics.

Example: check readiness

curl -s http://127.0.0.1:8091/v1/health
data-store --config config.toml --health

Ingestion monitor

data-store --config config.toml --monitor

The read-only terminal dashboard refreshes every 200 ms on one screen: ingestion stages and queues, committed annotation coverage, retrieval publication, grouped model calls, timings, reported token usage, waits, failures, and recent outcomes. Progress bars use measured totals; files and deduplicated active known sources remain separate. Blue, amber, and magenta accompany text state labels. Stale connections and display overflow are explicit; resize for more space. Q, Esc, or Ctrl-C exits. There are no monitor configuration settings or submenus. See SPEC-CLIENT.md §1.7 for polling and terminal behavior.

The headline distinguishes service readiness from ingestion completion and explains blocked work. Annotation and publication percentages cover active documents only. Persisted parse and queue failures remain visible after restart; routine checks leave idle panels unchanged.

To unblock failed work after addressing its cause:

data-store --config config.toml --clear-failures

This preserves successful work and failure history, resets retry eligibility, and resumes background processing. Command success does not mean ingestion is complete. Held parses and validation checks remain in force; workers terminated by panic or failed startup still require a restart.

Web UI

--serve <host>:<port> is a startup mode of the same binary, not a verb: it runs a local HTTP server in the foreground until the process is terminated, printing the bound URL. --config still applies — serve mode reuses the client's service base URL and admin token file.

data-store --config config.toml --serve 127.0.0.1:8092

The UI offers a query console (constraints, passage limit, evidence toggles, debug) showing cited passages with raw evidence and diagnostics in details panels, a unit explorer with relationship filters, a source view, a health dashboard, sync status, the held-parses list, an operation viewer, and the vocabulary explorer.

v1 is read-only. The browser reaches the service only through a fixed allowlist of proxy routes under /api/*, which exposes no mutating route; admin reads use the client's admin token file, and the UI itself carries no authentication.

Configuration

Configuration is a single TOML file (config.toml, from config.example.toml). Missing required settings and unknown keys are fatal startup errors. Configuration comments specify units, enforcement scope, and excess-input behavior. The sections:

Section Purpose
[server] HTTP bind address, required nonnegative integer startup_delay_seconds (0 disables the wait), and request-shape limits.
[logging] File-backed service logging.
[admin] Admin token file location.
[client] Client deadlines, polling cadence, and provenance previews.
[retrieval], [retrieval.entity_matching] Result counts, candidate depths, graph matching, passage/evidence budgets, and query admission.
[indexing] ColBERT-token caps for fine chunks, ColBERT windows, and context windows; context windows per excerpt; minimum fill ratio.
[resources] Read, allocation, and inventory admission guards.
[workers] Annotation/projection batching, concurrency, polling, and publication allowances.
[sqlite] Lock-wait timeout, cooperative SQL execution timeout, and progress-callback cadence.
[parsing] Candidate unit, relationship, and warning caps and the unit-body byte limit.
[scheduling] Sync cadence, backoff, and maintenance polling.
[diagnostics] Operational error, identifier, and progress-summary bounds.
[inference] Accelerator selection for local retrieval models. Required but unused when dense, ColBERT, and the reranker all use HTTP; no local accelerator is initialized in that mode.
[storage] Corpus root and service-owned index root.
[connectors.filesystem] Governance domain stamped on acquired sources.
[epub] EPUB admission budgets, all required: archive member count, per-member and total decompressed bytes, XML document bytes, image bytes, and element nesting depth.
[models] Dense, ColBERT, and reranker each select an exclusive local or http backend. Remote ColBERT uses vLLM /pooling token inference, a matching local tokenizer, persisted document matrices, and CPU MaxSim; no local ColBERT weights are loaded. The annotator uses an external chat-completions endpoint. See INSTALL.md for backend fields.
[policies] Paths to entity-match enable flags and annotator naming rules. Numeric entity-matching limits live in config.toml.

See INSTALL.md for the annotated example and the required absolute paths.

Model serving capacity is separate from grain size. models.dense.max_tokens and models.reranker.max_tokens declare the served capacities, and models.colbert.document_max_tokens the ColBERT document limit; startup requires /v1/models to advertise the configured capacity, requires indexing.colbert_max_tokens to equal the ColBERT document limit, and requires a context window plus its section-path prefix budget to fit the dense and reranker capacities. Server-side truncation is disabled; oversized HTTP inputs fail visibly. A ColBERT window whose prefixed input would exceed the document limit is embedded without the prefix.

Construction settings are recorded with new projections. Restore validates their recorded settings or explicit legacy formats without inference. Current resource guards may refuse large historical artifacts without declaring them corrupt. [parsing] and [epub] admission limits do not change parser identities.

The EPUB worker parses .epub files (EPUB 2 and EPUB 3) in-process with no external tool. It records only structure the source declares: sections from the navigation document or NCX and from headings, print pages from declared page-break markers, lists, tables decomposed to cells, figures with their image bytes archived by hash, captions, code blocks from pre, asides, and footnote and cross-reference links. No structure is inferred from text content; a source that declares no semantics for a feature gets no such feature. DRM-protected archives are recorded parse failures. Other formats are converted to plain text outside this service; only text/plain and EPUB are routed.

Annotation settings

These [models.annotator] settings are required; config.example.toml holds the shipped values. Excerpt size is not an annotator setting: it follows from indexing.context_max_tokens and indexing.excerpt_windows.

Setting Meaning
timeout_seconds Deadline for each model call, including thinking.
max_completion_tokens Completion-token allowance per call, including reasoning.
annotation_max_retries Malformed-output retries after the initial attempt.
annotation_retry_interval_seconds Fixed malformed-output retry interval, with no backoff ceiling.
execution_max_retries Execution-failure retries after the initial attempt.
execution_retry_initial_delay_seconds Initial execution-failure retry delay.
execution_retry_max_delay_seconds Execution-failure backoff ceiling.

The two failure counters are independent. Execution failures include timeouts, HTTP/protocol errors, token-limit termination, and internal producer failures. Their delay doubles up to the configured ceiling; it is independent of the configured scheduler waits and dense HTTP retry backoff. Malformed outputs use their fixed interval. Both limits allow 0 to disable that category's retries. Delays must be positive, and the execution ceiling must be at least the initial delay.

Temperature starts at 0.0 and becomes min(malformed_output_failures / annotation_max_retries, 1.0) on retries; execution failures do not advance it. With a zero annotation retry allowance, no temperature division or malformed-output retry occurs.

Once either counter exceeds its allowance, the annotation remains failed and is skipped as exhausted. Cancellation and scheduling deferrals spend neither budget. Counters and eligibility timers reset on restart or rebuild; failed chains restart as a whole, without intermediate-stage checkpoints.

About

Autonomous canonical content-graph and retrieval fabric for high-quality near real-time retrieval over rapidly changing corpora at scale with auditable provenance

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Contributors

Languages