Manage: persistent ingest history — every run (UI, CLI, kubectl) with running/stale states (#15) - #93
Merged
Merged
Conversation
… with running and stale states (#15) The Manage tab's jobs panel only knew jobs started from the UI (in-memory JobManager): a 'kubectl exec tpk ingest' -- the way production is re-indexed -- was invisible while it ran and afterwards, and nothing survived a restart. ingest_repo now writes a 'started' row to kg_ingest_log when a run begins and its ok/failed twin (same run_id) when it ends; a new 'error' column (added idempotently on startup, backfilled '') carries the failure reason. ingest.list_runs folds the two rows per run, newest first, and reports a start without an end as 'running' (< 2h) or 'stale' (the process died mid-run). ingest.latest_status applies the same rule to the corpus table's Last-ingest cell. GET /api/ingest-log?limit=&entry= (corpus:view) exposes it; the Manage tab's panel now lists that history, with a UI-driven job still overlaid live (phase, progress) until its end row lands. No migration: append-only stream, additive column, existing rows keep working (they just have no start time / duration). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #15.
Problem
The Manage tab's "Ingest jobs" panel only reflected the in-memory
JobManager: akubectl exec … tpk ingest— how production was re-indexed for 0.0.6 — was invisible both while running and afterwards, and nothing survived a restart. The persistent record (kg_ingest_log) already existed but only fed the corpus table's "Last ingest" cell, and only recorded finished runs with no error text.Change
startedrow when a run begins and itsok/failedtwin (samerun_id) when it ends. A newerrorcolumn carries the failure reason (RuntimeError: graphify exploded: exit 137). Logging the start is best-effort — it can never stop an ingest.ingest.list_runsfolds the two rows per run, newest first. A start with no end isrunning(< 2 h) orstale— the process died mid-run (OOM-kill, pod restart) and will never write the end row.ingest.latest_statusapplies the same rule to the corpus table, so "Last ingest" can't show "started" forever.GET /api/ingest-log?limit=50&entry=(corpus:view, like/api/repos;limit1–200):{runs: [{run_id, entry_key, status, nodes, edges, git_sha, error, started_at, finished_at, duration_s}]}. Rows written before this change have no start row →started_at/duration_sarenull.Schema / upgrade
No migration and nothing breaking: the stream is append-only, the
errorcolumn is added with_add_column_if_missingon startup (idempotent, backfilled"", same mechanism asdaily_token_limit), and existing rows keep working. Rolling back leaves inert extra rows/column.GET /api/jobsunchanged; export/import unaffected (the column has a default).Testing
startedthenokwith the samerun_id, and the start row is visible during extraction; a failure writesfailedwith the error.runningvsstale;entryfilter,limitand its bounds; the corpus table reports an abandoned start asstale.pytest: 343 passed, 10 skipped, 1 failed —tests/test_ingest.py::test_ingest_repo_passes_extraction_and_backend, environment-dependent, failing onmaintoo. Web build clean.tpk servein a throwaway database seeded the way real runs write: all four states rendered (running / failed with error / ok with duration / stale), corpus table showsstartedfor the live run andstalefor the abandoned one.Not in scope: cancelling a run, log retention, and per-phase progress for CLI runs (they write only start + end).
🤖 Generated with Claude Code