Skip to content

Commit bea26a7

Browse files
Replace text inside a draft field (t34)
Replace text inside a draft field
2 parents 12b5d82 + 165e253 commit bea26a7

8 files changed

Lines changed: 232 additions & 5 deletions

File tree

‎CHANGELOG.md‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -8,4 +8,5 @@
88
- `in2lambda source add FILE` freezes a document: it converts .docx and .tex to markdown beside the file, and writes a `draft.json` holding the markdown's hash and every block in it with the lines it spans, so that another tool can quote the source by line range. `in2lambda source show` prints that markdown numbered with the block ids. Freezing a file that has changed since is refused unless `--start-over` says to discard the draft, and so is showing one, since its block ids would name lines they are not the ids of. Both need pandoc and the `convert` extra, as `convert` does.
99
- A draft now holds a `log` of every command that changed it and a `fields` map of what those commands wrote, each field recording which layer wrote it (1 a spec, 2 a predicate, 3 a line range, 4 a literal), the source ranges it was copied from, whether it has been edited and by whom. `in2lambda draft mark ignore BLOCK` is the first such command, and `in2lambda draft replay` rebuilds the draft from the frozen markdown and the log, refusing unless what it builds is the `draft.json` that is there, byte for byte. A `draft.json` written before this has no `log` in it and is refused as one nothing here wrote; `in2lambda source add --start-over` freezes the document again.
1010
- A draft is filled in by `in2lambda draft question add`, `in2lambda draft part add QUESTION` and `in2lambda draft question solution QUESTION`. Each takes `--text` to copy the wording out of the frozen source, as a block id such as `b3` or as lines such as `s10:14`, or `--literal TEXT` where the source does not say it in a form the field can take, which records the field as edited and written by layer 4 rather than 3. Question and part numbers are worked out from the fields already written rather than given, so a replay arrives at the same ids. `in2lambda draft split block BLOCK AT` cuts a block the parser made one of two things into `b3a` and `b3b`, so that each half can be quoted on its own. A command writing a field that is already written, or quoting lines another field was taken from, is refused naming both fields.
11+
- `in2lambda draft field replace FIELD OLD NEW` changes the wording inside a field that is already written, for the faults only an edit can fix - a brace the OCR dropped out of some maths, which no range of the source says correctly. OLD has to occur in the field exactly once, or the command is refused saying how many times it occurs; `--regex` reads it as a regular expression and NEW as what to replace it with. The field is left quoting the lines it was taken from, at the layer that wrote it, but recorded as edited and by whoever replaced the wording, so the change can be shown against the source.
1112
- The Python API is unchanged: `in2lambda.main.runner` and everything under `in2lambda.api` take the same arguments and return the same objects.

‎in2lambda/draft/__init__.py‎

Lines changed: 76 additions & 5 deletions
Original file line numberDiff line numberDiff line change
@@ -68,6 +68,14 @@ class AlreadyFilled(SourceError):
6868
"""A command would write a field that is written, or lines another field took."""
6969

7070

71+
class NoSuchField(SourceError):
72+
"""A command changes the wording of a field the draft has not got as text."""
73+
74+
75+
class NotOnce(SourceError):
76+
"""The wording a command replaces is not in the field exactly once."""
77+
78+
7179
class ReplayDiffers(SourceError):
7280
"""Replaying a draft's log does not reproduce the draft."""
7381

@@ -120,13 +128,15 @@ def record(
120128
121129
Raises:
122130
AlreadyFilled: the field is written already, or the lines it was to be copied
123-
from are where another field came from. Nothing here changes a field once it
124-
is written, so either is a mistake, and worth naming both halves of.
131+
from are where another field came from. Nothing here writes a field twice -
132+
`field replace` changes the wording of one rather than writing it again -
133+
so either is a mistake, and worth naming both halves of.
125134
"""
126135
if key in draft["fields"]:
127136
raise AlreadyFilled(
128-
f"{key} is already written, and no command here changes a field that is. "
129-
"Run in2lambda source add --start-over to begin the draft again."
137+
f"{key} is already written, and no command here writes a field twice. Run "
138+
"in2lambda draft field replace to change the wording it holds, or "
139+
"in2lambda source add --start-over to begin the draft again."
130140
)
131141
for filled, field in draft["fields"].items():
132142
for taken in field["ranges"]:
@@ -177,7 +187,7 @@ def _argument(args: dict[str, Any], name: str, command: str, kind: type = str) -
177187
f'"{name}" argument.'
178188
)
179189
if not isinstance(args[name], kind):
180-
wanted = "a line number" if kind is int else "a name"
190+
wanted = {int: "a line number", bool: "true or false"}.get(kind, "a name")
181191
raise MalformedCommand(
182192
f"{args!r} in the log is not a command {command} can run: its "
183193
f'"{name}" is {args[name]!r} rather than {wanted}.'
@@ -473,6 +483,67 @@ def _question_solution(
473483
)
474484

475485

486+
@command("field replace")
487+
def _field_replace(
488+
draft: dict[str, Any], markdown: str, args: dict[str, Any], by: str
489+
) -> str:
490+
"""Replaces one piece of wording inside a field that is written already.
491+
492+
Some faults can only be fixed by changing the text: a brace the OCR dropped leaves
493+
maths KaTeX will not render, and no range of the source says it correctly. The
494+
layer and the ranges are left as they were, so the change can still be shown
495+
against the lines the field was taken from, and `edited` says what is there now is
496+
not what those lines say.
497+
498+
Raises:
499+
MalformedCommand: an argument is missing, or ``old`` is not a regular
500+
expression with ``regex``.
501+
NoSuchField: the draft has no field of that name holding text.
502+
NotOnce: ``old`` is not in the field exactly once, so which of it was meant is
503+
not something to guess at.
504+
"""
505+
key = _argument(args, "field", "field replace")
506+
old = _argument(args, "old", "field replace")
507+
new = _argument(args, "new", "field replace")
508+
# Only when it is there, so that a command nobody passed --regex to records no
509+
# argument for it, as every other option of a command does.
510+
regex = "regex" in args and _argument(args, "regex", "field replace", bool)
511+
512+
field = draft["fields"].get(key)
513+
if field is None or not isinstance(field.get("value"), str):
514+
raise NoSuchField(
515+
f"There is no field {key} holding text in {DRAFT}: field replace changes "
516+
"the wording of a field one of the commands before it has written."
517+
)
518+
value = field["value"]
519+
try:
520+
found = len(re.findall(old, value)) if regex else value.count(old)
521+
# A function rather than new itself, because re.sub reads a string as a
522+
# template, in which \t is a tab and \frac is an error. What is being repaired
523+
# here is LaTeX, so NEW is what gets written, backslashes and all.
524+
replaced = (
525+
re.sub(old, lambda _: new, value, count=1)
526+
if regex
527+
else value.replace(old, new, 1)
528+
)
529+
except re.error as error:
530+
raise MalformedCommand(
531+
f"{args!r} in the log is not a command field replace can run: it is not a "
532+
f"regular expression - {error}."
533+
) from None
534+
if found != 1:
535+
raise NotOnce(
536+
f"{old!r} occurs {found} times in {key} rather than once, so there is no "
537+
"one place in it to replace. Give more of the wording around it, or pass "
538+
"--regex and a pattern that matches it alone."
539+
)
540+
541+
field["value"] = replaced
542+
field["edited"] = True
543+
field["by"] = by
544+
return key
545+
546+
476547
@command("split block")
477548
def _split_block(
478549
draft: dict[str, Any], markdown: str, args: dict[str, Any], by: str

‎in2lambda/main.py‎

Lines changed: 24 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -369,6 +369,30 @@ def draft_split_block(block: str, at: int, by: str) -> None:
369369
_run("split block", {"block": block, "at": at}, by)
370370

371371

372+
@draft_group.group("field")
373+
def draft_field() -> None:
374+
"""Changes the wording of a field the draft has written already."""
375+
376+
377+
@draft_field.command("replace")
378+
@click.argument("field")
379+
@click.argument("old")
380+
@click.argument("new")
381+
@click.option(
382+
"--regex",
383+
is_flag=True,
384+
help="Read OLD as a regular expression, and NEW as what to replace it with.",
385+
)
386+
@_by
387+
def draft_field_replace(field: str, old: str, new: str, regex: bool, by: str) -> None:
388+
"""Replaces OLD with NEW in FIELD, which OLD has to occur exactly once in."""
389+
_run(
390+
"field replace",
391+
{"field": field, "old": old, "new": new, "regex": True if regex else None},
392+
by,
393+
)
394+
395+
372396
@draft_group.command("replay")
373397
def draft_replay() -> None:
374398
"""Rebuilds the draft in this directory from its log and checks it is the same."""

‎tests/fixtures/drafts/README.md‎

Lines changed: 4 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -12,3 +12,7 @@ rubric, two questions with a part each and a separate solutions section, written
1212
command there is: the second question runs into its part with no blank line between them, so the
1313
parser makes one block of the two and `split block` cuts it, and the first question's part is typed
1414
out rather than quoted, because the source writes it with an `(a)` the field should not carry.
15+
`field_replace` is a sheet the OCR left a brace out of the maths of: one `field replace` puts the
16+
brace back and another, with `--regex`, writes a `\tfrac` over the division in the solution, so
17+
both fields end up edited while their ranges still name the lines they were quoted from, and the
18+
backslash in what the second one writes is written rather than read as a replacement template.
Lines changed: 36 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,36 @@
1+
[
2+
{
3+
"args": {
4+
"text": "b2"
5+
},
6+
"by": "tests",
7+
"command": "question add"
8+
},
9+
{
10+
"args": {
11+
"question": "q1",
12+
"text": "b4"
13+
},
14+
"by": "tests",
15+
"command": "question solution"
16+
},
17+
{
18+
"args": {
19+
"field": "q1.text",
20+
"new": "\\mathrm{m/s}$",
21+
"old": "\\mathrm{m/s$"
22+
},
23+
"by": "ocr-fixer",
24+
"command": "field replace"
25+
},
26+
{
27+
"args": {
28+
"field": "q1.solution",
29+
"new": "\\tfrac{c}{f}",
30+
"old": "c\\s*/\\s*f",
31+
"regex": true
32+
},
33+
"by": "ocr-fixer",
34+
"command": "field replace"
35+
}
36+
]
Lines changed: 26 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,26 @@
1+
{
2+
"q1.solution": {
3+
"by": "ocr-fixer",
4+
"edited": true,
5+
"layer": 3,
6+
"ranges": [
7+
[
8+
8,
9+
8
10+
]
11+
],
12+
"value": "The wavelength is $\\lambda = \\tfrac{c}{f} = 0.78\\,\\mathrm{m}$."
13+
},
14+
"q1.text": {
15+
"by": "ocr-fixer",
16+
"edited": true,
17+
"layer": 3,
18+
"ranges": [
19+
[
20+
3,
21+
4
22+
]
23+
],
24+
"value": "A sound wave travels through still air at $c = 343\\,\\mathrm{m/s}$.\nFind the wavelength of a 440 Hz tone."
25+
}
26+
}
Lines changed: 8 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,8 @@
1+
# Sound in air
2+
3+
A sound wave travels through still air at $c = 343\,\mathrm{m/s$.
4+
Find the wavelength of a 440 Hz tone.
5+
6+
## Solutions
7+
8+
The wavelength is $\lambda = c / f = 0.78\,\mathrm{m}$.

‎tests/test_draft.py‎

Lines changed: 57 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -138,6 +138,27 @@ def test_a_log_naming_a_command_nothing_has_is_refused(
138138
({"command": "split block", "args": {"block": "b2"}, "by": "tests"}, "at"),
139139
({"command": "mark ignore", "args": {"block": 12}, "by": "tests"}, "12"),
140140
({"command": "question add", "args": {"text": 12}, "by": "tests"}, "12"),
141+
(
142+
{
143+
"command": "field replace",
144+
"args": {"field": "b1.ignore", "old": "a"},
145+
"by": "tests",
146+
},
147+
"new",
148+
),
149+
(
150+
{
151+
"command": "field replace",
152+
"args": {
153+
"field": "b1.ignore",
154+
"old": "a",
155+
"new": "b",
156+
"regex": "yes",
157+
},
158+
"by": "tests",
159+
},
160+
"true or false",
161+
),
141162
],
142163
ids=[
143164
"not an object",
@@ -151,6 +172,8 @@ def test_a_log_naming_a_command_nothing_has_is_refused(
151172
"no at argument",
152173
"a block that is a number",
153174
"a text that is a number",
175+
"no new argument",
176+
"a regex that is neither true nor false",
154177
],
155178
)
156179
def test_a_log_entry_that_is_not_a_command_is_refused(
@@ -278,6 +301,15 @@ def test_lines_another_field_was_taken_from_are_refused(
278301
(["draft", "part", "add", "q9", "--text", "s8"], "q9"),
279302
(["draft", "split", "block", "b3", "5"], "b3 is lines 5-6"),
280303
(["draft", "split", "block", "b3", "7"], "b3 is lines 5-6"),
304+
# q1.text has a $d$ and a $v$ in it, so a $ names four places and none of them.
305+
(["draft", "field", "replace", "q1.text", "$", "X"], "occurs 4 times"),
306+
(["draft", "field", "replace", "q1.text", "steam", "water"], "occurs 0 times"),
307+
(["draft", "field", "replace", "q9.text", "a", "b"], "q9.text"),
308+
(["draft", "field", "replace", "b1.ignore", "a", "b"], "b1.ignore"),
309+
(
310+
["draft", "field", "replace", "q1.text", "(", "X", "--regex"],
311+
"not a regular expression",
312+
),
281313
],
282314
ids=[
283315
"lines the source has not got",
@@ -288,6 +320,11 @@ def test_lines_another_field_was_taken_from_are_refused(
288320
"a question nothing has written",
289321
"a split at the line the block starts on",
290322
"a split past the line it ends on",
323+
"wording the field says more than once",
324+
"wording the field does not say",
325+
"a field nothing has written",
326+
"a field that is not text",
327+
"a regex that is not one",
291328
],
292329
)
293330
def test_a_command_naming_what_the_draft_has_not_got_is_refused(
@@ -326,6 +363,26 @@ def test_a_command_says_what_it_wrote(tmp_path: Path, monkeypatch) -> None:
326363
assert result.exit_code == 0, result.output
327364
assert result.output == "Wrote b3a and b3b.\n"
328365

366+
# And a replacement names the field it changed, which it leaves quoting the same
367+
# lines as before, said to be edited, and by whoever replaced the wording.
368+
result = runner.invoke(
369+
cli,
370+
["draft", "field", "replace", "q1.solution", "d^2", "d^{2}", "--by", "ocr"],
371+
)
372+
373+
assert result.exit_code == 0, result.output
374+
assert result.output == "Wrote q1.solution.\n"
375+
draft = json.loads(draft_path.read_text())
376+
assert draft["fields"]["q1.solution"] == {
377+
"value": "The flow rate is $Q = \\pi d^{2} v / 4$.",
378+
"layer": 3,
379+
"ranges": [[16, 16]],
380+
"edited": True,
381+
"by": "ocr",
382+
}
383+
# --regex is an option, so a command nobody passed it to logs no argument for it.
384+
assert "regex" not in draft["log"][-1]["args"]
385+
329386

330387
def test_the_halves_of_a_split_block_are_blocks_like_any_other(
331388
tmp_path: Path, monkeypatch

0 commit comments

Comments
 (0)