Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,4 +8,5 @@
- `in2lambda source add FILE` freezes a document: it converts .docx and .tex to markdown beside the file, and writes a `draft.json` holding the markdown's hash and every block in it with the lines it spans, so that another tool can quote the source by line range. `in2lambda source show` prints that markdown numbered with the block ids. Freezing a file that has changed since is refused unless `--start-over` says to discard the draft, and so is showing one, since its block ids would name lines they are not the ids of. Both need pandoc and the `convert` extra, as `convert` does.
- A draft now holds a `log` of every command that changed it and a `fields` map of what those commands wrote, each field recording which layer wrote it (1 a spec, 2 a predicate, 3 a line range, 4 a literal), the source ranges it was copied from, whether it has been edited and by whom. `in2lambda draft mark ignore BLOCK` is the first such command, and `in2lambda draft replay` rebuilds the draft from the frozen markdown and the log, refusing unless what it builds is the `draft.json` that is there, byte for byte. A `draft.json` written before this has no `log` in it and is refused as one nothing here wrote; `in2lambda source add --start-over` freezes the document again.
- A draft is filled in by `in2lambda draft question add`, `in2lambda draft part add QUESTION` and `in2lambda draft question solution QUESTION`. Each takes `--text` to copy the wording out of the frozen source, as a block id such as `b3` or as lines such as `s10:14`, or `--literal TEXT` where the source does not say it in a form the field can take, which records the field as edited and written by layer 4 rather than 3. Question and part numbers are worked out from the fields already written rather than given, so a replay arrives at the same ids. `in2lambda draft split block BLOCK AT` cuts a block the parser made one of two things into `b3a` and `b3b`, so that each half can be quoted on its own. A command writing a field that is already written, or quoting lines another field was taken from, is refused naming both fields.
- `in2lambda draft field replace FIELD OLD NEW` changes the wording inside a field that is already written, for the faults only an edit can fix - a brace the OCR dropped out of some maths, which no range of the source says correctly. OLD has to occur in the field exactly once, or the command is refused saying how many times it occurs; `--regex` reads it as a regular expression and NEW as what to replace it with. The field is left quoting the lines it was taken from, at the layer that wrote it, but recorded as edited and by whoever replaced the wording, so the change can be shown against the source.
- The Python API is unchanged: `in2lambda.main.runner` and everything under `in2lambda.api` take the same arguments and return the same objects.
81 changes: 76 additions & 5 deletions in2lambda/draft/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -68,6 +68,14 @@ class AlreadyFilled(SourceError):
"""A command would write a field that is written, or lines another field took."""


class NoSuchField(SourceError):
"""A command changes the wording of a field the draft has not got as text."""


class NotOnce(SourceError):
"""The wording a command replaces is not in the field exactly once."""


class ReplayDiffers(SourceError):
"""Replaying a draft's log does not reproduce the draft."""

Expand Down Expand Up @@ -120,13 +128,15 @@ def record(

Raises:
AlreadyFilled: the field is written already, or the lines it was to be copied
from are where another field came from. Nothing here changes a field once it
is written, so either is a mistake, and worth naming both halves of.
from are where another field came from. Nothing here writes a field twice -
`field replace` changes the wording of one rather than writing it again -
so either is a mistake, and worth naming both halves of.
"""
if key in draft["fields"]:
raise AlreadyFilled(
f"{key} is already written, and no command here changes a field that is. "
"Run in2lambda source add --start-over to begin the draft again."
f"{key} is already written, and no command here writes a field twice. Run "
"in2lambda draft field replace to change the wording it holds, or "
"in2lambda source add --start-over to begin the draft again."
)
for filled, field in draft["fields"].items():
for taken in field["ranges"]:
Expand Down Expand Up @@ -177,7 +187,7 @@ def _argument(args: dict[str, Any], name: str, command: str, kind: type = str) -
f'"{name}" argument.'
)
if not isinstance(args[name], kind):
wanted = "a line number" if kind is int else "a name"
wanted = {int: "a line number", bool: "true or false"}.get(kind, "a name")
raise MalformedCommand(
f"{args!r} in the log is not a command {command} can run: its "
f'"{name}" is {args[name]!r} rather than {wanted}.'
Expand Down Expand Up @@ -473,6 +483,67 @@ def _question_solution(
)


@command("field replace")
def _field_replace(
draft: dict[str, Any], markdown: str, args: dict[str, Any], by: str
) -> str:
"""Replaces one piece of wording inside a field that is written already.

Some faults can only be fixed by changing the text: a brace the OCR dropped leaves
maths KaTeX will not render, and no range of the source says it correctly. The
layer and the ranges are left as they were, so the change can still be shown
against the lines the field was taken from, and `edited` says what is there now is
not what those lines say.

Raises:
MalformedCommand: an argument is missing, or ``old`` is not a regular
expression with ``regex``.
NoSuchField: the draft has no field of that name holding text.
NotOnce: ``old`` is not in the field exactly once, so which of it was meant is
not something to guess at.
"""
key = _argument(args, "field", "field replace")
old = _argument(args, "old", "field replace")
new = _argument(args, "new", "field replace")
# Only when it is there, so that a command nobody passed --regex to records no
# argument for it, as every other option of a command does.
regex = "regex" in args and _argument(args, "regex", "field replace", bool)

field = draft["fields"].get(key)
if field is None or not isinstance(field.get("value"), str):
raise NoSuchField(
f"There is no field {key} holding text in {DRAFT}: field replace changes "
"the wording of a field one of the commands before it has written."
)
value = field["value"]
try:
found = len(re.findall(old, value)) if regex else value.count(old)
# A function rather than new itself, because re.sub reads a string as a
# template, in which \t is a tab and \frac is an error. What is being repaired
# here is LaTeX, so NEW is what gets written, backslashes and all.
replaced = (
re.sub(old, lambda _: new, value, count=1)
if regex
else value.replace(old, new, 1)
)
except re.error as error:
raise MalformedCommand(
f"{args!r} in the log is not a command field replace can run: it is not a "
f"regular expression - {error}."
) from None
if found != 1:
raise NotOnce(
f"{old!r} occurs {found} times in {key} rather than once, so there is no "
"one place in it to replace. Give more of the wording around it, or pass "
"--regex and a pattern that matches it alone."
)

field["value"] = replaced
field["edited"] = True
field["by"] = by
return key


@command("split block")
def _split_block(
draft: dict[str, Any], markdown: str, args: dict[str, Any], by: str
Expand Down
24 changes: 24 additions & 0 deletions in2lambda/main.py
Original file line number Diff line number Diff line change
Expand Up @@ -369,6 +369,30 @@ def draft_split_block(block: str, at: int, by: str) -> None:
_run("split block", {"block": block, "at": at}, by)


@draft_group.group("field")
def draft_field() -> None:
"""Changes the wording of a field the draft has written already."""


@draft_field.command("replace")
@click.argument("field")
@click.argument("old")
@click.argument("new")
@click.option(
"--regex",
is_flag=True,
help="Read OLD as a regular expression, and NEW as what to replace it with.",
)
@_by
def draft_field_replace(field: str, old: str, new: str, regex: bool, by: str) -> None:
"""Replaces OLD with NEW in FIELD, which OLD has to occur exactly once in."""
_run(
"field replace",
{"field": field, "old": old, "new": new, "regex": True if regex else None},
by,
)


@draft_group.command("replay")
def draft_replay() -> None:
"""Rebuilds the draft in this directory from its log and checks it is the same."""
Expand Down
4 changes: 4 additions & 0 deletions tests/fixtures/drafts/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -12,3 +12,7 @@ rubric, two questions with a part each and a separate solutions section, written
command there is: the second question runs into its part with no blank line between them, so the
parser makes one block of the two and `split block` cuts it, and the first question's part is typed
out rather than quoted, because the source writes it with an `(a)` the field should not carry.
`field_replace` is a sheet the OCR left a brace out of the maths of: one `field replace` puts the
brace back and another, with `--regex`, writes a `\tfrac` over the division in the solution, so
both fields end up edited while their ranges still name the lines they were quoted from, and the
backslash in what the second one writes is written rather than read as a replacement template.
36 changes: 36 additions & 0 deletions tests/fixtures/drafts/field_replace/commands.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,36 @@
[
{
"args": {
"text": "b2"
},
"by": "tests",
"command": "question add"
},
{
"args": {
"question": "q1",
"text": "b4"
},
"by": "tests",
"command": "question solution"
},
{
"args": {
"field": "q1.text",
"new": "\\mathrm{m/s}$",
"old": "\\mathrm{m/s$"
},
"by": "ocr-fixer",
"command": "field replace"
},
{
"args": {
"field": "q1.solution",
"new": "\\tfrac{c}{f}",
"old": "c\\s*/\\s*f",
"regex": true
},
"by": "ocr-fixer",
"command": "field replace"
}
]
26 changes: 26 additions & 0 deletions tests/fixtures/drafts/field_replace/expected.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,26 @@
{
"q1.solution": {
"by": "ocr-fixer",
"edited": true,
"layer": 3,
"ranges": [
[
8,
8
]
],
"value": "The wavelength is $\\lambda = \\tfrac{c}{f} = 0.78\\,\\mathrm{m}$."
},
"q1.text": {
"by": "ocr-fixer",
"edited": true,
"layer": 3,
"ranges": [
[
3,
4
]
],
"value": "A sound wave travels through still air at $c = 343\\,\\mathrm{m/s}$.\nFind the wavelength of a 440 Hz tone."
}
}
8 changes: 8 additions & 0 deletions tests/fixtures/drafts/field_replace/source.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,8 @@
# Sound in air

A sound wave travels through still air at $c = 343\,\mathrm{m/s$.
Find the wavelength of a 440 Hz tone.

## Solutions

The wavelength is $\lambda = c / f = 0.78\,\mathrm{m}$.
57 changes: 57 additions & 0 deletions tests/test_draft.py
Original file line number Diff line number Diff line change
Expand Up @@ -138,6 +138,27 @@ def test_a_log_naming_a_command_nothing_has_is_refused(
({"command": "split block", "args": {"block": "b2"}, "by": "tests"}, "at"),
({"command": "mark ignore", "args": {"block": 12}, "by": "tests"}, "12"),
({"command": "question add", "args": {"text": 12}, "by": "tests"}, "12"),
(
{
"command": "field replace",
"args": {"field": "b1.ignore", "old": "a"},
"by": "tests",
},
"new",
),
(
{
"command": "field replace",
"args": {
"field": "b1.ignore",
"old": "a",
"new": "b",
"regex": "yes",
},
"by": "tests",
},
"true or false",
),
],
ids=[
"not an object",
Expand All @@ -151,6 +172,8 @@ def test_a_log_naming_a_command_nothing_has_is_refused(
"no at argument",
"a block that is a number",
"a text that is a number",
"no new argument",
"a regex that is neither true nor false",
],
)
def test_a_log_entry_that_is_not_a_command_is_refused(
Expand Down Expand Up @@ -278,6 +301,15 @@ def test_lines_another_field_was_taken_from_are_refused(
(["draft", "part", "add", "q9", "--text", "s8"], "q9"),
(["draft", "split", "block", "b3", "5"], "b3 is lines 5-6"),
(["draft", "split", "block", "b3", "7"], "b3 is lines 5-6"),
# q1.text has a $d$ and a $v$ in it, so a $ names four places and none of them.
(["draft", "field", "replace", "q1.text", "$", "X"], "occurs 4 times"),
(["draft", "field", "replace", "q1.text", "steam", "water"], "occurs 0 times"),
(["draft", "field", "replace", "q9.text", "a", "b"], "q9.text"),
(["draft", "field", "replace", "b1.ignore", "a", "b"], "b1.ignore"),
(
["draft", "field", "replace", "q1.text", "(", "X", "--regex"],
"not a regular expression",
),
],
ids=[
"lines the source has not got",
Expand All @@ -288,6 +320,11 @@ def test_lines_another_field_was_taken_from_are_refused(
"a question nothing has written",
"a split at the line the block starts on",
"a split past the line it ends on",
"wording the field says more than once",
"wording the field does not say",
"a field nothing has written",
"a field that is not text",
"a regex that is not one",
],
)
def test_a_command_naming_what_the_draft_has_not_got_is_refused(
Expand Down Expand Up @@ -326,6 +363,26 @@ def test_a_command_says_what_it_wrote(tmp_path: Path, monkeypatch) -> None:
assert result.exit_code == 0, result.output
assert result.output == "Wrote b3a and b3b.\n"

# And a replacement names the field it changed, which it leaves quoting the same
# lines as before, said to be edited, and by whoever replaced the wording.
result = runner.invoke(
cli,
["draft", "field", "replace", "q1.solution", "d^2", "d^{2}", "--by", "ocr"],
)

assert result.exit_code == 0, result.output
assert result.output == "Wrote q1.solution.\n"
draft = json.loads(draft_path.read_text())
assert draft["fields"]["q1.solution"] == {
"value": "The flow rate is $Q = \\pi d^{2} v / 4$.",
"layer": 3,
"ranges": [[16, 16]],
"edited": True,
"by": "ocr",
}
# --regex is an option, so a command nobody passed it to logs no argument for it.
assert "regex" not in draft["log"][-1]["args"]


def test_the_halves_of_a_split_block_are_blocks_like_any_other(
tmp_path: Path, monkeypatch
Expand Down
Loading