Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,7 @@

## Unreleased — initial research scaffold

- Add eight fictional challenge cards, a contract-question map, a case template, and a partner discovery guide. Keep their unresolved semantics outside the unchanged `0.1.0` executable contract.
- Establish the organisational thesis, claims ledger, and baseline-first research plan.
- Propose a draft `0.1.0` commitment event export with version lineage and occurrence/availability timestamps.
- Add a wholly fictional trajectory, context projection utility, and temporal/semantic validation tests.
Expand Down
4 changes: 4 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -68,6 +68,8 @@ A simple model that wins is a successful research result. A shared model valid f
| [Foundations](docs/foundations.md) | The organisation, human agency, and learning from priors |
| [Research plan](docs/research-plan.md) | Hypotheses, comparisons, evaluation, and stop conditions |
| [Data contract](docs/data-contract.md) | What to collect, label, connect, and keep out |
| [Fictional challenge set](examples/challenge-set/README.md) | Eight difficult or contrasting cases that question the contract and its scope |
| [Partner discovery guide](docs/partner-discovery.md) | A first conversation that lets partners introduce their own distinctions |
| [Application integration](docs/application-integration.md) | How apps contribute to the Organisation Core |
| [First 90 days](docs/first-90-days.md) | A bounded starting project and collaboration questions |
| [Reading list](docs/reading-list.md) | Intellectual lineage and the limits of the evidence |
Expand All @@ -86,6 +88,8 @@ python3 -m venv .venv

This validates the draft event contract and constructs the context available at a historical prediction cutoff. It demonstrates exclusion of late-arriving evidence, future actions, and outcome assessments from predictor inputs. It does **not** train a model or establish that a dataset is suitable for research. See [the example walkthrough](examples/README.md).

For discussion before collecting data, use the [eight fictional challenge cards](examples/challenge-set/README.md). They preserve unresolved interpretations and deliberately include work that may not fit the commitment abstraction. They are authored scenarios, not labelled training data; the current schema remains unchanged.

The validation command reports **11 valid fictional events**. The Tuesday-noon snapshot contains **e01, e02, and e03**: the accepted promise, a reported delay, and a proposed response. An observation made earlier but received Wednesday is correctly excluded. The test suite checks these boundaries and the preservation of both commitment versions.

## Project status
Expand Down
2 changes: 2 additions & 0 deletions docs/first-90-days.md
Original file line number Diff line number Diff line change
Expand Up @@ -13,6 +13,8 @@ Sequence/graph models and JEPA follow only if these results justify them. Long c

## Initial work packages

The [fictional challenge set](../examples/challenge-set/README.md) and [partner discovery guide](partner-discovery.md) provide starting material for the first two work packages. Gather participants' own accounts before showing the cards; no case constitutes a partner finding or a training label.

1. **Target definition:** choose a recurring promise, specify fulfilment and observation, and identify where revisions make labels ambiguous.
2. **Contract review:** challenge the action vocabulary, version lineage, beneficiary role, late evidence, and missing-data semantics with fictional counterexamples.
3. **Measurement adapter:** export one prospective trajectory across two surfaces, preserving authority and provenance.
Expand Down
53 changes: 53 additions & 0 deletions docs/partner-discovery.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,53 @@
# Partner discovery with fictional cases

Purpose: discover whether the problem, unit of analysis, and outcome definitions fit a partner's work before proposing a dataset. This is a conversation guide, not a validated interview instrument or a request for data access.

The [challenge set](../examples/challenge-set/README.md) contains eight wholly fictional discussion cards. It is deliberately small and incomplete. No partner findings have been collected through this guide yet.

## A first conversation: about 40 minutes

**First 10 minutes — their work before our vocabulary.** Ask: “Can you describe a recurring piece of work where several people depend on what happens next?” Invite a fictional or suitably general account. Let them name the roles, purpose, and expectations. Follow with “How do you know it went well?” and “What happens when circumstances change?” Avoid introducing commitment, failure, JEPA, or our action categories as required answers.

**Next 10 minutes — establish one decision point.** Ask what was known at that moment, what was missing, who could act, and where evidence would appear later. Distinguish the date something happened from when the relevant person or system learned of it. Ask what useful work never enters the records and what would be burdensome to capture. A participant should not need to disclose documents or names.

**Next 15 minutes — use two or three cards.** Choose cases that illuminate the participant's account, including a routine or poor-fit case where relevant. Show the decision point first; keep the later account and maintainer discussion collapsed. Ask for an interpretation, then reveal what happened and ask what changed. Do not use the maintainer notes to mark answers correct. Ask which details feel implausible, which cases are missing, and whether the underlying abstraction helps.

**Final 5 minutes — agree one small follow-up.** It could be a new fictional case, a target-definition review, or a diagram of where evidence lives. Agree what may be retained and shared. There is no requirement to offer data, endorse the project, or join a pilot.

## Preserve disagreement and the source of an idea

If several people attend, invite independent initial interpretations before group discussion when practical. Record differences rather than forcing consensus. Distinguish an idea the participant introduced unaided from an answer prompted by a card or by the facilitator. The latter may be useful, but it is weaker evidence that the issue was salient without prompting.

Use these prompts sparingly:

- “Who would describe this differently?”
- “What evidence would change that interpretation?”
- “What can the existing systems already do well?”
- “Is this a data-model problem, a working-practice problem, or something we should leave outside the project?”
- “What would make this effort cost more than it helps?”

## A lightweight note template

This is a blank template, not an example of a completed interview. Keep working notes outside the public repository. Retain only what the participant agrees to; do not record the session by default.

| Field | What to record |
| --- | --- |
| Agreed purpose and handling | What may be noted, retained, abstracted, and shared; no implied research/data permission |
| Work in the participant's terms | Their unit of work and desired result, before using project vocabulary |
| Unaided observations | Issues and distinctions raised before showing any card |
| Cutoff and visibility | What information existed and when the relevant people/system could access it |
| Cards and order shown | Stable IDs and whether the later account or maintainer notes had been revealed |
| Interpretations | Distinct accounts, including unresolved disagreements and rejection of the framing |
| Missing evidence and burden | What would distinguish the accounts and the cost of observing it |
| Candidate consequence | A question, target, procedure, contract change, or exclusion; mark as tentative |
| Follow-up | A bounded artefact, willing owner if agreed, and what may be made public |

Do not copy personal identifiers, confidential examples, private quotes, or organisational records into public issues. Permission to have a conversation is not permission to publish its contents or train on them. Follow the repository's [data stewardship proposal](data-contract.md#data-stewardship).

## Turn discovery into a reviewable change

Start with a finding candidate, not a generator: describe the abstract ambiguity, plausible interpretations, and evidence needed. Keep a single anecdote distinct from a recurring pattern; record contrary accounts and unknowns. Do not report interview counts as prevalence in organisations or assert saturation from this small exercise.

If an issue can be illustrated safely, write a new fictional card and ask whether it preserves the important distinction. A public issue can then propose one of four outcomes: clarify the target; improve observation/adjudication; extend the contract with tests; or exclude the case from the current prediction task. Leave the choice open when evidence is insufficient.

Bulk synthetic generation becomes useful when a reviewed question calls for controlled variation or engineering scale tests. It should encode declared, contestable assumptions. It cannot supply evidence that those assumptions describe organisations or that a model trained on them will transfer to partner data.
2 changes: 2 additions & 0 deletions examples/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,8 @@

Every entity, event, source, and probability in `fictional-trajectory.jsonl` is invented. This is an executable specification example, not a training dataset or benchmark. No model generated the example's probability.

For cases that challenge the contract and its scope, see the [fictional challenge set](challenge-set/README.md). Those discussion cards include unresolved semantics and are not JSONL fixtures. Start partner conversations with the [discovery guide](../docs/partner-discovery.md).

The fictional team promises an integration by Friday 9 January. It later agrees a Monday 12 January deadline with the customer. Friday's original promise is assessed as not met; Monday's revised promise is assessed as met. Both assessments remain attached to their own versions.

| Event | Meaning |
Expand Down
33 changes: 33 additions & 0 deletions examples/challenge-set/C01-delivered-but-rejected.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,33 @@
# C01 — Delivered on time, rejected by the beneficiary

Wholly fictional discussion case. No canonical outcome label. Facilitators should reveal the later account only after discussing the decision point.

## At the decision point

A service team accepts a promise to provide a customer with a weekly export by Friday at 17:00. The recorded criteria specify the file format, required columns, and delivery location. The customer intends to use the export in an existing reconciliation process; no representative sample or compatibility check was agreed.

On Thursday at 12:00, the delivery owner knows the file passes the recorded format checks. The customer's operations lead has asked whether it will work in their process, but that question has no recorded answer. The owner must decide whether to investigate compatibility before delivery.

**Before revealing the outcome:** What has actually been promised? What evidence is available for a forecast, and whose interpretation matters?

<details>
<summary>Reveal the later account</summary>

The file arrives Friday at 16:30 and passes every recorded check. On Monday the customer rejects it: the values use a convention their reconciliation process cannot interpret. Delivery records completion; customer operations records that the expected benefit was not received. Nobody disputes when the file arrived.

</details>

## Questions for the participant

- Would you call this fulfilment, failure, or a badly formed promise? What would distinguish those accounts?
- Does the beneficiary's rejection identify an existing obligation or introduce a new requirement?
- What small clarification at the start would have prevented the ambiguity, and who could authorise it?

<details>
<summary>Maintainer discussion — provisional, not an answer key</summary>

Under a narrow interpretation, the written criteria were met and the realised value was poor. Under a broader accepted understanding, compatibility was part of the promise and fulfilment is disputed. Preserve both accounts until the scope is resolved; do not silently rewrite the original criteria or make satisfaction automatically determine every label.

The draft can store criteria, observations, and a disputed assessment. It lacks structured beneficiary acceptance and realised-value dimensions. A candidate discovery is whether these require separate targets, better acceptance procedures, or both. Ask for the agreed scope, a sample agreed before the deadline, and evidence of who could accept it. This case supplies none of that missing evidence.

</details>
35 changes: 35 additions & 0 deletions examples/challenge-set/C02-responsible-cancellation.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,35 @@
# C02 — A responsible cancellation

Wholly fictional discussion case. No canonical outcome label.

## At the decision point

A facilities team accepts a promise to equip a training room by Day 20. The authorised scope includes furniture, equipment, and a working demonstration. By Day 8, equipment has been ordered but installation has not begun.

On Day 9, the organisation decides to leave the building. The requester and facilities owner know this. They can continue installation, seek permission to stop, or negotiate a different use for the equipment. Some purchases are refundable; others have already incurred costs.

**Before revealing the outcome:** Which obligation still matters, and who may release the team from it?

<details>
<summary>Reveal the later account</summary>

On Day 10, the authorised sponsor and requester agree to cancel the room setup. The team returns refundable equipment and documents the remaining costs. No room is equipped by Day 20. The requester reports that a usable room in the old building would now have no value. The scenario does not establish exactly what would have been spent had work continued.

</details>

## Questions for the participant

- Which record distinguishes a legitimate release from a team quietly abandoning a promise?
- What obligations survive cancellation, such as handling costs or informing dependent teams?
- How would you compare responsible stopping with unnecessary cancellation without rewarding either automatically?

<details>
<summary>Maintainer discussion — provisional, not an answer key</summary>

The original physical deliverable was not produced. That fact alone does not establish a poor organisational decision. A completion-only target and a decision-quality evaluation answer different questions. Cancellation also does not erase the original promise or guarantee the change was wise.

The draft can describe a `revise_stop` action but does not provide a ratified withdrawal state or a release-of-obligations contract. An abandoned action is not a cancelled commitment. Rewriting the intent of a later version may hide rather than solve this distinction.

Keep this trajectory outside the initial binary benchmark until its target policy is explicit. Candidate needs include authority, beneficiary agreement, residual obligations, and observed costs. Claims about avoided costs would require additional assumptions or a separate comparison. A partner might instead recommend excluding this commitment family altogether.

</details>
35 changes: 35 additions & 0 deletions examples/challenge-set/C03-competing-commitments.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,35 @@
# C03 — Two promises, one shared resource

Wholly fictional discussion case. No canonical outcome label.

## At the decision point

Two teams accept separate customer onboarding promises, A and B, both due Friday. Each requires one full day of the same specialist before the customer can verify the result. The specialist has one day available that week. Neither customer agreed to a priority rule.

At Tuesday noon, each team can see its own promise and a reference to the specialist. Both believe a day is reserved. The organisation's allocation record, available to a coordinator, shows only one reservation slot. That discrepancy has not reached either delivery team. An organisational predictor would need a declared access scope: what one team knows is not automatically what the prediction system knows.

**Before revealing the outcome:** What would you need to know to forecast these promises together? Who can resolve the conflict?

<details>
<summary>Reveal the later account</summary>

On Wednesday, a coordinator discovers the double reservation. Under an explicit escalation mandate, the organisation allocates the specialist to A and agrees a Monday revision for B with its customer. A is fulfilled Friday. B misses its original deadline and is fulfilled Monday. The customer affected by B incurs disruption that the scenario does not quantify.

</details>

## Questions for the participant

- Where should evidence of the shared constraint have appeared, and when was it available to whom?
- Does A's success demonstrate a better team or simply the allocation decision?
- What would make the priority decision acceptable to both beneficiaries?

<details>
<summary>Maintainer discussion — provisional, not an answer key</summary>

Separate labels can describe fulfilment of the respective versions, but independent feature rows may conceal the shared cause. A resource allocation might improve one prediction target while worsening another. This case cannot identify the best priority policy.

The draft supports multiple commitments and opaque dependency references. It does not structure capacity, reservation conflicts, access scope, or priority authority. Free-text descriptions would not automatically enable graph inference about contention.

Candidate evidence includes capacity with units and effective dates, reservation history, each promise's consequences, and the authority to trade them off. Related trajectories may need to stay together when designing evaluation splits. Discover whether the partner's records can support that grouping before treating each promise as an independent example.

</details>
37 changes: 37 additions & 0 deletions examples/challenge-set/C04-late-evidence.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,37 @@
# C04 — Late evidence changes the assessment

Wholly fictional discussion case. No canonical outcome label.

## At the decision point

A storage service commits to keeping a shipment within an agreed temperature range through Friday at 17:00. Its acceptance criteria identify the monitoring interval and the measurement source.

At Thursday noon, all readings available to the prediction system are within range. A sensor has recorded additional readings locally, but those readings have not been transmitted. There is no basis at that cutoff to assert what Friday's measurements will show.

**Before revealing the outcome:** Which clocks must be preserved so a future training run can reproduce Thursday's information?

<details>
<summary>Reveal the later account</summary>

Friday at 17:20, the monitoring service reports a breach during the agreed interval. An assessor records `not_met` with that evidence. On Tuesday, a reviewed correction shows the transmitted series used an incorrect conversion. The corrected source readings support a later `met` assessment against the same original criteria. The original report and both assessments remain in the history.

The corrected readings concern Friday, but their corrected interpretation was first available Tuesday. Moving that knowledge back to Friday would change what the earlier predictor could have known.

</details>

## Questions for the participant

- What evidence permits the second assessment to supersede the first?
- What label would a training run on Monday have had access to? What changes for a later run?
- Would repeated evidence corrections reveal a delivery problem, a measurement problem, or both?

<details>
<summary>Maintainer discussion — provisional, not an answer key</summary>

The draft's occurrence and availability timestamps support late evidence; the context utility excludes it before availability. The draft also allows multiple assessments of one version. It has no typed correction link, adjudication rule, or label-selection implementation.

Selecting the last event by timestamp is not a sufficient governance rule. A later assertion can also be wrong or unauthorised. Preserve assessment provenance and distinguish a historical replay using labels then available from an evaluation using subsequently adjudicated outcomes. Both can be useful, but must be declared separately.

Candidate evidence includes the correction's basis, authorised adjudication, and immutable dataset manifests. This story does not establish that any real measurement source has those properties.

</details>
Loading