Skip to content

Emit a versioned JSON manifest of the resolved run plan for external tools #1120

Description

@jpn--

Request

Please emit a documented, versioned JSON manifest describing the resolved execution plan for a particular ActivitySim run, for both single-process and multiprocess execution. A suggested filename is run_manifest.json, alongside the existing log artifacts. The essential request is a supported data contract, rather than particular field names.

Motivation

Lighthouse (the under development Boston model) has an external CPU/memory and model-progress monitor that currently hard-codes model names, their order, multiprocess phase names, and completion-message patterns. Adding a custom component or splitting a phase therefore requires an unrelated monitor edit. Otherwise, components can become invisible or progress tracking can stall. This surfaced in CTPSSTAFF/lighthouse#21.

ActivitySim already resolves much of this information in get_run_list() in activitysim/core/mp_tasks.py and writes run_list.txt. However, print_run_list() explicitly describes that output as informational. Its repeated phase blocks and Python-style values are not a documented JSON/YAML interchange format.

Maintaining another model list duplicates the source of truth. Reading raw settings in each external tool duplicates inheritance, overrides, process-count defaults, and resume logic. Discovering steps solely from completion logs cannot reveal future steps or reliably establish expected workers: parallel completion order is not model order, and a monitor attaching late may miss worker-start messages.

Proposed minimum contract

Generate the manifest from ActivitySim's resolved execution state after relevant configuration and command-line overrides have been applied.

Information What consumers need
Schema identity/version A version independent of the ActivitySim package version, with documented compatibility rules.
Invocation identity An identifier unique to this execution attempt, UTC generation time, and ActivitySim version. Reuse existing run identity facilities where appropriate, distinguishing a resumed invocation from its predecessor.
Ordered model steps Full resolved sequence, including extension models, with unambiguous step identifiers/ordinals and execution names for correlating logs. Preserve invocation arguments/modifiers where applicable.
Ordered phases Actual phase names and step membership, without assumptions such as mp_households. Represent single-process execution explicitly under the same contract.
Expected workers Resolved worker counts and identifiers/log process names per phase, including one-worker phases. Distinguish the logical worker set from workers that must launch on this attempt.
Resume semantics Requested resume value and effective position after checkpoint/breadcrumb resolution, including work already satisfied and work remaining. A literal "_" alone is insufficient. Multiprocess resumes may require per-worker positions or explicit unknown values.
Artifact discovery Documented manifest location and references to relevant log/checkpoint artifacts where available, with clear relative-path rules.

Use ordered arrays and JSON-native values, including null, with documented meanings for missing/unknown values. Consumers should not infer phase membership from model names or order from dictionary keys.

Some resume details may only become known when workers restore checkpoints. Please document that boundary rather than treating a requested checkpoint as an already-resolved plan. An initial implementation could expose those details as unknown, with a later structured update mechanism.

Publication and compatibility

  • Publish once the plan is valid and as early as practical before model execution, allowing monitors to attach at startup or later.
  • Write atomically so readers cannot observe partially written JSON.
  • Distinguish fresh/resumed execution attempts so reused output directories do not silently associate stale plans with new runs.
  • Retain run_list.txt for human inspection and existing consumers.
  • Document the schema, ideally with JSON Schema and example manifests. Allow additive fields; change the schema version for incompatible changes.
  • Export focused execution metadata rather than arbitrary settings, environment variables, or input records.
  • Document supported CLI/programmatic entry points and provide a shared export path where feasible.

Scope: execution plan versus live status

This request is for an execution-plan artifact, not a new scheduler or complete telemetry service. Planned work must not be presented as successful completion. A monitor still needs execution evidence, and a fraction of steps completed is not an estimate of elapsed-time progress.

A separate follow-up could provide structured start/completion/failure events carrying the same invocation, phase, step, and worker identifiers. Keeping events separate would allow the manifest to land without replacing existing logging.

Other potential consumers

These are potential uses, not claims of existing integrations:

  • Terminal/web progress dashboards: discover custom components, phase boundaries, and expected workers without model-specific code.
  • Resource profiling and benchmarking tools: join CPU/memory samples and timings to steps; compare runs with different sequences and process layouts.
  • CI and regression tooling: check that the resolved plan contains an extension step, detect ordering/partition changes, and distinguish full from resumed/partial runs. Actual execution still requires separate evidence.
  • HPC/cloud job wrappers and workflow orchestrators: display intended work and concurrency, associate artifacts with the correct attempt, and explain remaining work on restart without reimplementing settings resolution.
  • Run archives, provenance reports, and support diagnostics: preserve the applicable execution plan even after settings files change. This complements rather than replaces full reproducibility metadata.
  • Notebook/report generators: enumerate components and phases and organize timing/output information without scraping presentation-oriented text.

Suggested acceptance coverage

Small fixtures should verify that the manifest matches the plan actually used for single-process/multiprocess runs; custom model and phase names; inherited settings and command-line overrides; one-worker/multiworker phases; explicit checkpoint resumes and resume_after: _; and partial worker completion before restart.

Adding/reordering a model or splitting a phase should change the manifest automatically, with no external monitor code change.

Would a supported export of the resolved run plan be an appropriate upstream interface? These fields are a proposal; the key requirement is that external consumers can obtain the actual plan through a stable, versioned contract.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions