diff --git a/CHANGELOG.md b/CHANGELOG.md index 74beef5..9d0de84 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -9,4 +9,5 @@ - A draft now holds a `log` of every command that changed it and a `fields` map of what those commands wrote, each field recording which layer wrote it (1 a spec, 2 a predicate, 3 a line range, 4 a literal), the source ranges it was copied from, whether it has been edited and by whom. `in2lambda draft mark ignore BLOCK` is the first such command, and `in2lambda draft replay` rebuilds the draft from the frozen markdown and the log, refusing unless what it builds is the `draft.json` that is there, byte for byte. A `draft.json` written before this has no `log` in it and is refused as one nothing here wrote; `in2lambda source add --start-over` freezes the document again. - A draft is filled in by `in2lambda draft question add`, `in2lambda draft part add QUESTION` and `in2lambda draft question solution QUESTION`. Each takes `--text` to copy the wording out of the frozen source, as a block id such as `b3` or as lines such as `s10:14`, or `--literal TEXT` where the source does not say it in a form the field can take, which records the field as edited and written by layer 4 rather than 3. Question and part numbers are worked out from the fields already written rather than given, so a replay arrives at the same ids. `in2lambda draft split block BLOCK AT` cuts a block the parser made one of two things into `b3a` and `b3b`, so that each half can be quoted on its own. A command writing a field that is already written, or quoting lines another field was taken from, is refused naming both fields. - `in2lambda draft field replace FIELD OLD NEW` changes the wording inside a field that is already written, for the faults only an edit can fix - a brace the OCR dropped out of some maths, which no range of the source says correctly. OLD has to occur in the field exactly once, or the command is refused saying how many times it occurs; `--regex` reads it as a regular expression and NEW as what to replace it with. The field is left quoting the lines it was taken from, at the layer that wrote it, but recorded as edited and by whoever replaced the wording, so the change can be shown against the source. +- `in2lambda spec run SPEC` runs a YAML file of selectors over the frozen source: it says which blocks are questions, parts and solutions, which to ignore, what to strip off the front of each one, and which of the four filters lays the solutions out. It fills in the draft's fields with the markdown of the lines each was taken from, records the spec's name and hash in the log so a replay runs the same file, and reports every block it made nothing of. Running an edited spec over a draft it has already filled in is refused, as freezing a document that has changed is: `in2lambda source add --start-over` begins the draft again. Reading a spec needs pyyaml, which the `convert` extra now installs alongside panflute. See [the spec page](https://lambda-feedback.github.io/in2lambda/spec.html) for the selectors and layouts. - The Python API is unchanged: `in2lambda.main.runner` and everything under `in2lambda.api` take the same arguments and return the same objects. diff --git a/docs/source/index.md b/docs/source/index.md index 54342d2..816852c 100644 --- a/docs/source/index.md +++ b/docs/source/index.md @@ -63,6 +63,7 @@ A fully type-annotated extensively documented Python library is available for th quickstart question-format filters/index +spec ``` ```{toctree} diff --git a/docs/source/spec.md b/docs/source/spec.md new file mode 100644 index 0000000..3a4f7bb --- /dev/null +++ b/docs/source/spec.md @@ -0,0 +1,103 @@ +# 📐 Specs + +A spec is a small YAML file saying which blocks of a document are questions, which are parts and +which are solutions. Running one fills in the draft beside the document, so that the wording of +every question comes out of the source rather than being retyped: + +```bash +$ in2lambda source add questions.docx +$ in2lambda spec run spec.yaml +b6 is in no field. +``` + +The last line is the point of it: a spec run reports every block it made nothing of, so what is +left to account for is in front of you rather than quietly missing. + +The fields a draft holds belong to the spec that wrote them, so a spec is run over a draft once. +Running an edited one again is refused; freeze the document afresh and run it, which is two +commands: + +```bash +$ in2lambda source add --start-over questions.docx +$ in2lambda spec run spec.yaml +``` + +## What a spec says + +```yaml +question: Header level=2 text~'^Question' +part: ListItem +solution: after Header text=Solutions, label~'^\d+(\([a-z]\))?$' +strip: ['^#+ ', '^\([a-z]\) ', '^\d+(\([a-z]\))? '] +ignore: Header level=1 +layout: PartsSepSol +``` + +`question` and `layout` have to be there; `part`, `solution`, `strip` and `ignore` need not be. + +- **`question`, `part`, `solution`** select the blocks that are each of those things. +- **`ignore`** selects the blocks that are none of them - a running header, a page of + instructions - and marks them as `in2lambda draft mark ignore` would, so they are not reported + as left out. +- **`strip`** is a list of patterns taken off the front of every value: the `(a) ` or `1. ` that + labels a part in the document, but not in the question. +- **`layout`** is one of the [filters](filters/index), and says which solution answers which + question or part. See below. + +A block is whatever the first of `ignore`, `question`, `part`, `solution` to match it says it is. +That order is fixed, whatever order the keys are written in, so a spec whose selectors overlap +has to tell them apart by what they match rather than by where they are in the file. + +## Selectors + +A selector is a block type, then any number of constraints: + +``` +[after SELECTOR,] [Type] name=value name~'regex' ... +``` + +The type is a pandoc element - `Header`, `Para`, `ListItem` - and may be left out to match any +block. A constraint is about one of three things: + +| Attribute | What it is | +|-----------|------------| +| `level` | A heading's level: `level=2` is `##`. | +| `text` | The whole block as text, with the markup taken off. | +| `label` | The first word of that text, which is usually what numbers a question. | + +`=` asks for exactly that; `~` for a regular expression anywhere in it. `after SELECTOR,` says +the block has to come after the first block that selector matches, which is how the solutions at +the end of a problem sheet are told apart from the questions at the front. + +A regular expression goes in single quotes. YAML reads `\(` inside double quotes as an escape +and complains, and `'^\([a-z]\)'` is the same string without the argument. + +A selector matches what **pandoc** makes of the document, while a field holds the **markdown** of +the lines it came from. That is worth knowing in two places: a part written `(a) Find the load.` +is a `ListItem`, because pandoc reads `(a)` as a list marker, and `strip` still has to take the +`(a) ` off the front of the value, because the line it was copied from still has it. + +## Layouts + +The layout is the one thing that differs between problem sheets that are otherwise alike: where +the solutions are, and what each of them answers. + +| Layout | Which solution answers what | +|--------|-----------------------------| +| `PartsOneSol` | One solution to the whole question, however many parts it has. | +| `PartSolPartSol` | Each solution answers the part just before it, or the question if it has no parts yet. | +| `PartPartSolSol` | The parts come together and their solutions come after, in the same order. | +| `PartsSepSol` | Every solution is at the end: the first answers the first part of the first question, and so on. | + +## What it writes + +Each question is `q1`, `q2` and so on in the order they appear, and each of its parts `q1.a`, +`q1.b`. So a spec fills in `q1.text`, `q1.a.text`, `q1.a.solution` and, for a question answered +as a whole, `q1.solution`. Every one of them records the lines it was copied from, and that a +spec wrote it. + +The spec is recorded in the draft's log with its hash, so `in2lambda draft replay` rebuilds the +same draft from the same spec - and refuses if the spec has been edited since, because then it +would be checking the draft against something else. That is why running an edited spec over a +draft it has already filled in is refused too: the draft would be left holding fields no spec on +disk wrote, and no replay could ever check it again. diff --git a/in2lambda/draft/__init__.py b/in2lambda/draft/__init__.py index d648219..cd88416 100644 --- a/in2lambda/draft/__init__.py +++ b/in2lambda/draft/__init__.py @@ -15,9 +15,12 @@ from pathlib import Path from typing import Any +import in2lambda.spec from in2lambda.source import ( DRAFT, SourceError, + _digest, + _elements, _require_conversion_tools, blocks, frozen, @@ -28,13 +31,15 @@ Command = dict[str, Any] """One entry of the log: ``{"command": name, "args": {...}, "by": who}``.""" -Handler = Callable[[dict[str, Any], str, dict[str, Any], str], str] -"""What a command does: `handler(draft, markdown, args, by)`, changing the draft. +Handler = Callable[[dict[str, Any], str, dict[str, Any], str, str], str] +"""What a command does: `handler(draft, markdown, args, by, directory)`. The frozen markdown is passed in rather than read, so that a handler quoting the source -by line range quotes the same text on a replay as it did when it first ran. What comes -back is what the command wrote, named - the key of the field, or the block ids a split -made - which is what whoever ran it needs in the command after this one. +by line range quotes the same text on a replay as it did when it first ran; the +directory is where the draft is, which is what a file a command names is beside. What +comes back is what the command wrote, named - the key of the field, or the block ids a +split made - which is what whoever ran it needs in the command after this one, or what +it left out where a command wrote a draft's worth of fields at once. """ _HANDLERS: dict[str, Handler] = {} @@ -80,6 +85,10 @@ class ReplayDiffers(SourceError): """Replaying a draft's log does not reproduce the draft.""" +class SpecChanged(SourceError): + """The spec file a log names is not the one that ran: it has changed, or gone.""" + + def command(name: str) -> Callable[[Handler], Handler]: """Registers a handler as the command of that name. @@ -169,6 +178,24 @@ def _fault(entry: Any) -> str: return "" +def _checked(entry: Any) -> Command: + """One entry of a log, given that it has the shape of a command. + + Both readers of a log come through here - the one applying an entry and the one + looking over the entries already applied - so that a hand-edited log says the same + thing whichever of them reads it first. + + Raises: + MalformedCommand: the entry is not a command. + """ + if fault := _fault(entry): + raise MalformedCommand( + f"{entry!r} in the log is not a command: it {fault}. A command is an " + 'object with a "command" naming it, its "args", and who it was run "by".' + ) + return entry + + def _argument(args: dict[str, Any], name: str, command: str, kind: type = str) -> Any: """One argument of a command, given that the log entry gave it as `kind`. @@ -195,7 +222,9 @@ def _argument(args: dict[str, Any], name: str, command: str, kind: type = str) - return args[name] -def apply(draft: dict[str, Any], markdown: str, entry: Any) -> str: +def apply( + draft: dict[str, Any], markdown: str, entry: Any, directory: str = "." +) -> str: """Runs one command against a draft and records it in the draft's log. Args: @@ -204,6 +233,7 @@ def apply(draft: dict[str, Any], markdown: str, entry: Any) -> str: entry: The command, as it is written in the log. Anything at all, rather than a `Command`, because a log is read from a file anyone can edit: what shape it has is something to tell the reader about, not something to assume. + directory: Where the draft is, and so what a file the command names is beside. Returns: What the command wrote, as the handler names it. @@ -212,17 +242,13 @@ def apply(draft: dict[str, Any], markdown: str, entry: Any) -> str: MalformedCommand: the entry is not a command. UnknownCommand: nothing is registered under that name. """ - if fault := _fault(entry): - raise MalformedCommand( - f"{entry!r} in the log is not a command: it {fault}. A command is an " - 'object with a "command" naming it, its "args", and who it was run "by".' - ) + entry = _checked(entry) if (handler := _HANDLERS.get(entry["command"])) is None: raise UnknownCommand( f"{entry['command']} is not a command this version of in2lambda has, so " "the draft cannot be built from its log. It was written by a newer one." ) - written = handler(draft, markdown, entry["args"], entry["by"]) + written = handler(draft, markdown, entry["args"], entry["by"], directory) # After the handler, so a command that was refused is not recorded as having run. draft["log"].append(entry) return written @@ -236,15 +262,15 @@ def execute(entry: Command, directory: str = ".") -> str: directory: Where the ``draft.json`` to change is. Returns: - What the command wrote, as the handler names it: the key of a field, or the - block ids a split made. + What the command wrote, as the handler names it: the key of a field, the block + ids a split made, or what a spec run left in no field at all. Raises: SourceError: the draft is missing, is not one of ours, or was written from markdown that has changed since; or the command is unknown or refused. """ draft, markdown = frozen(directory) - written = apply(draft, markdown, entry) + written = apply(draft, markdown, entry, directory) save(Path(directory) / DRAFT, draft) return written @@ -264,6 +290,7 @@ def replay(directory: str = ".") -> None: the commands would be replayed against lines they were not run against. MalformedCommand: the log holds something that is not a command. UnknownCommand: the log names a command nothing here registered. + SpecChanged: a spec the log was run with has changed or gone since. ReplayDiffers: the rebuilt draft is not the one on disk, byte for byte. """ _require_conversion_tools() @@ -278,7 +305,7 @@ def replay(directory: str = ".") -> None: "fields": {}, } for entry in draft["log"]: - apply(rebuilt, markdown, entry) + apply(rebuilt, markdown, entry, directory) path = Path(directory) / DRAFT if serialise(rebuilt) != path.read_bytes(): @@ -289,6 +316,32 @@ def replay(directory: str = ".") -> None: ) +def coverage(draft: dict[str, Any]) -> list[str]: + """The blocks of a draft that nothing has made anything of yet. + + Args: + draft: The draft to look over. + + Returns: + The ids of the blocks that are in no field and have not been ignored, in + document order. A spec run prints these: they are what is left to account for, + and an empty list is the whole document spoken for. + """ + ranges = [ + line_range + for field in draft["fields"].values() + for line_range in field["ranges"] + ] + return [ + block["id"] + for block in draft["blocks"] + if f"{block['id']}.ignore" not in draft["fields"] + and not any( + start <= block["end"] and block["start"] <= end for start, end in ranges + ) + ] + + def _block(draft: dict[str, Any], block: str) -> dict[str, Any]: """One block of the frozen source, given the draft has one of that id. @@ -419,7 +472,7 @@ def _require_question(draft: dict[str, Any], question: str, command: str) -> Non @command("mark ignore") def _mark_ignore( - draft: dict[str, Any], markdown: str, args: dict[str, Any], by: str + draft: dict[str, Any], markdown: str, args: dict[str, Any], by: str, directory: str ) -> str: """Marks one block of the frozen source as nothing to take a question from.""" block = _argument(args, "block", "mark ignore") @@ -436,7 +489,7 @@ def _mark_ignore( @command("question add") def _question_add( - draft: dict[str, Any], markdown: str, args: dict[str, Any], by: str + draft: dict[str, Any], markdown: str, args: dict[str, Any], by: str, directory: str ) -> str: """Adds a question, taking the first number no question has taken.""" return _fill( @@ -451,7 +504,7 @@ def _question_add( @command("part add") def _part_add( - draft: dict[str, Any], markdown: str, args: dict[str, Any], by: str + draft: dict[str, Any], markdown: str, args: dict[str, Any], by: str, directory: str ) -> str: """Adds a part to a question, taking the first number that question has not.""" question = _argument(args, "question", "part add") @@ -468,7 +521,7 @@ def _part_add( @command("question solution") def _question_solution( - draft: dict[str, Any], markdown: str, args: dict[str, Any], by: str + draft: dict[str, Any], markdown: str, args: dict[str, Any], by: str, directory: str ) -> str: """Gives a question its worked solution, wherever in the source it is written.""" question = _argument(args, "question", "question solution") @@ -485,7 +538,7 @@ def _question_solution( @command("field replace") def _field_replace( - draft: dict[str, Any], markdown: str, args: dict[str, Any], by: str + draft: dict[str, Any], markdown: str, args: dict[str, Any], by: str, directory: str ) -> str: """Replaces one piece of wording inside a field that is written already. @@ -546,7 +599,7 @@ def _field_replace( @command("split block") def _split_block( - draft: dict[str, Any], markdown: str, args: dict[str, Any], by: str + draft: dict[str, Any], markdown: str, args: dict[str, Any], by: str, directory: str ) -> str: """Cuts one block of the frozen source in two, so each half can be named. @@ -570,3 +623,72 @@ def _split_block( {**found, "id": f"{block}b", "start": at}, ] return f"{block}a and {block}b" + + +def _spec_as_run(directory: str, name: str, digest: str) -> bytes: + """A spec the log says has run, given it is still there and still says the same. + + Raises: + SpecChanged: there is no such file beside the draft, or it is not the one the + log records running. Either way the fields it wrote are fields nothing on + disk would write again, so neither a replay nor another run can check them. + """ + try: + raw = (Path(directory) / name).read_bytes() + except FileNotFoundError: + raise SpecChanged( + f"There is no {name} beside {DRAFT}, and the log says the draft was filled " + "in from one. Put it back, or start the draft again with in2lambda source " + "add --start-over." + ) from None + if _digest(raw) != digest: + raise SpecChanged( + f"{name} has changed since it was run against {DRAFT}, so the fields it " + "wrote are not the ones it would write now. Put it back, or start the " + "draft again with in2lambda source add --start-over." + ) + return raw + + +@command("spec run") +def _spec_run( + draft: dict[str, Any], markdown: str, args: dict[str, Any], by: str, directory: str +) -> str: + """Fills in a draft's fields from a spec of selectors over the frozen source.""" + _require_conversion_tools() + name = _argument(args, "spec", "spec run") + digest = _argument(args, "hash", "spec run") + # Every spec the log says has run, rather than one named the same way as this one: a + # spec that has been edited since leaves fields the log can no longer reproduce + # whatever it is spelled as now, so what has run is what to check. On a replay this + # re-reads specs whose own entries checked them, which is a file read each. + for entry in map(_checked, draft["log"]): + if entry["command"] == "spec run": + _spec_as_run( + directory, + _argument(entry["args"], "spec", "spec run"), + _argument(entry["args"], "hash", "spec run"), + ) + raw = _spec_as_run(directory, name, digest) + + spec = in2lambda.spec.load(raw) + # The blocks the selectors run over are the ones the parser makes of the source, and + # a `split block` since has left the draft holding halves the parser never made. So + # an ignored block is named and ranged from here rather than from the draft: the + # field then spans the whole of what was ignored, and coverage, which goes by lines + # as well as by name, counts each half of a split block as covered by it. + elements = _elements(markdown) + fields, ignored = in2lambda.spec.fields(spec, elements, markdown) + for found in fields: + record(draft, found.key, found.value, layer=1, ranges=found.ranges, by=by) + lines = {block.id: [block.start, block.end] for block, _ in elements} + # The field `mark ignore` writes, so that coverage need not care which said so. + for block_id in ignored: + record( + draft, f"{block_id}.ignore", True, layer=1, ranges=[lines[block_id]], by=by + ) + # A spec writes a draft's worth of fields, so what it hands back is the other way + # round: what it made nothing of, which is what is left for anyone to act on. + if uncovered := coverage(draft): + return "\n".join(f"{block} is in no field." for block in uncovered) + return "Every block is in a field or ignored." diff --git a/in2lambda/main.py b/in2lambda/main.py index b4eacfa..0d163dd 100644 --- a/in2lambda/main.py +++ b/in2lambda/main.py @@ -15,6 +15,7 @@ Iterator, ) from contextlib import contextmanager +from pathlib import Path from typing import Any, Optional import rich_click as click @@ -30,6 +31,7 @@ from in2lambda.source import ( # noqa: F401 # Re-exported, so not unused. ConversionToolsMissing, SourceError, + _digest, _pandoc, _require_conversion_tools, file_type, @@ -401,5 +403,28 @@ def draft_replay() -> None: click.echo("Replays as it stands.") +@cli.group("spec") +def spec_group() -> None: + """Runs a YAML spec of selectors over the frozen source in this directory.""" + + +@spec_group.command("run") +@click.argument("spec", type=click.Path(exists=True, dir_okay=False)) +@_by +def spec_run(spec: str, by: str) -> None: + """Fills the draft's fields in from SPEC, and says which blocks it left out.""" + with _message_not_traceback(): + # The hash goes in the log beside the file's name, so that a replay can tell + # whether it is running the spec that wrote the fields it is checking. + report = in2lambda.draft.execute( + { + "command": "spec run", + "args": {"spec": spec, "hash": _digest(Path(spec).read_bytes())}, + "by": by, + } + ) + click.echo(report) + + if __name__ == "__main__": cli() diff --git a/in2lambda/source/__init__.py b/in2lambda/source/__init__.py index bd2d6dc..af03095 100644 --- a/in2lambda/source/__init__.py +++ b/in2lambda/source/__init__.py @@ -103,8 +103,13 @@ def _require_conversion_tools() -> None: missing = [] if shutil.which("pandoc") is None: missing.append("pandoc (see https://pandoc.org/installing.html)") - if importlib.util.find_spec("panflute") is None: - missing.append("panflute (pip install 'in2lambda[convert]')") + # Both come from the one extra, so they are named together rather than twice over. + if absent := [ + package + for module, package in (("panflute", "panflute"), ("yaml", "pyyaml")) + if importlib.util.find_spec(module) is None + ]: + missing.append(f"{' and '.join(absent)} (pip install 'in2lambda[convert]')") if missing: raise ConversionToolsMissing( f"Converting documents needs {' and '.join(missing)}." @@ -328,6 +333,16 @@ def blocks(markdown: str) -> list[Block]: >>> blocks("# Title\n\nSome words.\n") [Block(id='b1', type='heading', start=1, end=1), Block(id='b2', type='paragraph', start=3, end=3)] """ + return [block for block, _ in _elements(markdown)] + + +def _elements(markdown: str) -> list[tuple[Block, Any]]: + """Every block of some markdown, each beside the panflute element it was taken from. + + A selector matches on what the element is - its type, its heading level, the text it + stringifies to - which the block alone does not say, so anything matching against + the source takes this and projects the blocks out of it, as :func:`blocks` does. + """ import panflute as pf document = pf.convert_text( @@ -338,18 +353,22 @@ def blocks(markdown: str) -> list[Block]: # paragraph, a definition list - pandoc reports the first as running on into the # second's first line, so no block is allowed to reach where the next one starts, # nor past the end of the document. - limits = [start - 1 for _, start, _ in found[1:]] + [len(markdown.splitlines())] + limits = [start - 1 for _, start, _, _ in found[1:]] + [len(markdown.splitlines())] return [ - Block(f"b{number}", kind, start, min(end, limit)) - for number, ((kind, start, end), limit) in enumerate(zip(found, limits), 1) + (Block(f"b{number}", kind, start, min(end, limit)), element) + for number, ((kind, start, end, element), limit) in enumerate( + zip(found, limits), 1 + ) ] def _spans(element, pf): # type: ignore[no-untyped-def] - """The ``(type, start, end)`` triples one top-level element accounts for. + """The ``(type, start, end, element)`` quadruples one top-level element accounts for. A list is several: the ticket asks for a list item, not a list, and an item spans - everything nested under it. + everything nested under it. The element given back is the one that block is, past + the Div `sourcepos` wraps it in, so that whatever matches on it matches on what an + author would call it. """ inner = _unwrapped(element, pf) if isinstance(inner, (pf.BulletList, pf.OrderedList)): @@ -357,11 +376,11 @@ def _spans(element, pf): # type: ignore[no-untyped-def] # no element to take a position from, so there is no range to give it and it # is left out rather than guessed at. return [ - ("list item", _range(item.content[0])[0], _range(item.content[-1])[1]) + ("list item", _range(item.content[0])[0], _range(item.content[-1])[1], item) for item in inner.content if len(item.content) ] - return [(_kind(inner, pf), *_range(element))] + return [(_kind(inner, pf), *_range(element), inner)] def _unwrapped(element, pf): # type: ignore[no-untyped-def] diff --git a/in2lambda/spec/__init__.py b/in2lambda/spec/__init__.py new file mode 100644 index 0000000..10b503b --- /dev/null +++ b/in2lambda/spec/__init__.py @@ -0,0 +1,439 @@ +r"""Reads a spec of selectors, and says what each block of a frozen source is. + +A spec is a small YAML file saying which blocks of a document are questions, which are +parts and which are solutions, and which layout they are written in:: + + question: Header level=2 text~'^Question' + part: Para text~'^\([a-z]\) ' + solution: after Header text=Solutions, label~'^\d+' + strip: ['^\([a-z]\) ', '^\d+ '] + ignore: Header level=1 + layout: PartsSepSol + +Nothing here decides what a question is: the selectors say which blocks are which, and +the layout says how a solution is paired up with the question or part it answers, which +is the one thing that differs between the filters in :mod:`in2lambda.filters` and is +copied from them here. What comes out is one field per question, part and solution, so +that a draft written by a spec says the same things as a draft written by hand. + +Reading a spec needs pyyaml, which only the ``convert`` extra installs; the command that +calls this checks for it first, along with pandoc and panflute. +""" + +import re +from dataclasses import dataclass, field +from typing import Any, NamedTuple, Optional + +from in2lambda.filters import builtin_filters +from in2lambda.source import Block, SourceError + +_KEYS = ("question", "part", "solution", "strip", "ignore", "layout") +"""Everything a spec may say. Anything else in one is a typo, and is refused as one.""" + +_ATTRIBUTES = ("level", "text", "label") +"""What a constraint can be about: a heading's level, a block's text, its first word.""" + +_ROLES = ("ignore", "question", "part", "solution") +"""The selectors a spec holds, in the order a block is tried against them. + +A block is whatever the first of them to match it says it is. The order is this one +whatever order a spec writes its keys in: ignore before the rest so that a page nobody +wants is out of the way, and question before part so that a question numbered like one +of its own parts is still the question. +""" + +_TOKEN = re.compile( + r"""\s*(?:(?P\w+)\s*(?P[=~])\s*""" + r"""(?:"(?P[^"]*)"|'(?P[^']*)'|(?P[^\s,]+))|(?P\w+))""" +) +"""One word of a selector: a ``name=value`` or ``name~regex`` constraint, or a type.""" + + +class BadSpec(SourceError): + """A spec cannot be read, whether as YAML or as selectors.""" + + +def _refuse(line: int, message: str) -> BadSpec: + """A refusal of a spec, which always ends by saying which line to go and look at.""" + return BadSpec(f"{message} See line {line} of the spec.") + + +@dataclass +class Constraint: + """One ``level=2`` or ``text~'^Question'`` of a selector.""" + + attribute: str + wanted: "re.Pattern[str] | str" + + def holds(self, value: Optional[str]) -> bool: + """Whether a block's attribute is what this asks for, given the block has one.""" + if value is None: + return False + if isinstance(self.wanted, str): + return value == self.wanted + return bool(self.wanted.search(value)) + + +@dataclass +class Selector: + """Which blocks of a document a spec is talking about.""" + + type: Optional[str] = None + constraints: list[Constraint] = field(default_factory=list) + after: Optional["Selector"] = None + + def matches(self, elements: list[Any], index: int, pf: Any) -> bool: + """Whether the block at `index` is one of these. + + Args: + elements: Every block of the document, as the panflute element it is. + index: Which of them to decide about. + pf: The panflute module, imported by the caller that has it. + + Returns: + True if the element is of this type, meets every constraint, and comes + after something the ``after`` selector matches. + """ + element = elements[index] + if self.after is not None and not any( + self.after.matches(elements, earlier, pf) for earlier in range(index) + ): + return False + if self.type is not None and type(element).__name__ != self.type: + return False + return all( + constraint.holds(_attribute(constraint.attribute, element, pf)) + for constraint in self.constraints + ) + + +@dataclass +class Spec: + """What a spec file says, once it has been read.""" + + question: Selector + layout: str + part: Optional[Selector] = None + solution: Optional[Selector] = None + ignore: Optional[Selector] = None + strip: "list[re.Pattern[str]]" = field(default_factory=list) + + +class Field(NamedTuple): + """One field a spec fills in: what it is called, what it says, where it came from.""" + + key: str + value: str + ranges: list[list[int]] + + +def _attribute(name: str, element: Any, pf: Any) -> Optional[str]: + """What a block says for one attribute, or None where it has not got one.""" + if name == "level": + return str(element.level) if isinstance(element, pf.Header) else None + text = pf.stringify(element).strip() + if name == "text": + return text + return words[0] if (words := text.split()) else None + + +def _split(text: str) -> tuple[str, Optional[str]]: + """A selector either side of its comma, which a quoted regex may hold its own of.""" + quote = "" + for position, character in enumerate(text): + if quote: + if character == quote: + quote = "" + elif character in "\"'": + quote = character + elif character == ",": + return text[:position], text[position + 1 :] + return text, None + + +def _clause(text: str, line: int, after: Optional[Selector] = None) -> Selector: + """One ``[Type] constraint*`` of a selector, given it says nothing else.""" + selector = Selector(after=after) + position = 0 + while position < len(text): + if (token := _TOKEN.match(text, position)) is None: + raise _refuse( + line, + f"{text[position:].strip()!r} is not something a selector says. A " + "selector is a block type and then any number of name=value or " + "name~'regex' constraints.", + ) + position = token.end() + if (name := token["name"]) is None: + if selector.type is not None or selector.constraints: + raise _refuse( + line, + "A selector names one block type, before its constraints, so " + f"{token['type']} is one word too many.", + ) + selector.type = _type(token["type"], line) + else: + if name not in _ATTRIBUTES: + raise _refuse( + line, + f"{name} is not something a block has: a constraint is about " + f"{', '.join(_ATTRIBUTES)}.", + ) + wanted = next( + value + for value in (token["double"], token["single"], token["bare"]) + if value is not None + ) + selector.constraints.append( + Constraint( + name, _pattern(wanted, line) if token["operator"] == "~" else wanted + ) + ) + return selector + + +def _type(name: str, line: int) -> str: + """A block type, given pandoc has one of that name.""" + import panflute as pf + + found = getattr(pf, name, None) + if not (isinstance(found, type) and issubclass(found, pf.Element)): + raise _refuse( + line, + f"{name} is not a pandoc element. A selector names one as pandoc does - " + "Header, Para, ListItem - or leaves the type out to match any block.", + ) + return name + + +def _pattern(regex: Any, line: int) -> "re.Pattern[str]": + """A regex, given it is one. A backslash in YAML wants single quotes around it.""" + try: + return re.compile(regex) + except (re.error, TypeError) as error: # TypeError: a strip list of numbers. + raise _refuse( + line, f"{regex!r} is not a regular expression: {error}." + ) from None + + +def _selector(text: Any, line: int) -> Selector: + """One selector of a spec, as its `after` clause and the rest.""" + if not isinstance(text, str): + raise _refuse(line, f"A selector is a line of text, which {text!r} is not.") + head, tail = _split(text.strip()) + if head.strip().startswith("after "): + after = _clause(head.strip()[len("after ") :], line) + return _clause(tail.strip() if tail else "", line, after) + if tail is not None: + raise _refuse( + line, + "A selector's comma separates its `after` clause from the rest, and " + f"{text!r} has no `after` in it.", + ) + return _clause(head.strip(), line) + + +def load(text: "str | bytes") -> Spec: + r"""Reads a spec, given that it says what a spec says. + + Args: + text: The contents of the spec file, as text or as the bytes it was read as. + The bytes are handed to YAML rather than decoded here, since YAML knows + which encoding a file is in from its byte order mark and refuses one it + cannot read the way it refuses anything else about a spec. + + Returns: + The spec, with its selectors parsed and its strip patterns compiled. + + Raises: + BadSpec: the text is not YAML, is in an encoding YAML cannot read, is not a + mapping, says something a spec does not, or holds a selector, pattern or + layout that cannot be read. Every one of them says which line to look at. + + Examples: + >>> from in2lambda.spec import load + >>> load("question: Header level=2\nlayout: PartsOneSol\n").layout + 'PartsOneSol' + """ + import yaml + + # The composed nodes carry the line each key is written on; the values come from + # safe_load, which builds them rather than leaving them as nodes to unpick. Both + # are read here, since a file that composes can still fail to be built - a tag + # nothing constructs, a key nothing can hash - and that is as much a fault in the + # spec as a quote left open. + try: + node = yaml.compose(text) + given = yaml.safe_load(text) + except yaml.YAMLError as error: + mark = getattr(error, "problem_mark", None) + raise _refuse( + 1 if mark is None else mark.line + 1, f"The spec is not YAML: {error}." + ) from None + + if not isinstance(node, yaml.MappingNode): + raise _refuse(1, "A spec is a mapping of question, part, solution and so on.") + lines = { + key.value: key.start_mark.line + 1 + for key, _ in node.value + if isinstance(key, yaml.ScalarNode) + } + + # By str, because a key someone has written need not be one: `1: Header` is YAML. + if unknown := sorted(set(given) - set(_KEYS), key=str): + raise _refuse( + lines.get(unknown[0], 1), + f"{unknown[0]} is not something a spec says. A spec says " + f"{', '.join(_KEYS)}.", + ) + if missing := [key for key in ("question", "layout") if key not in given]: + raise _refuse( + 1, + "A spec says which blocks are questions and how they are laid out, so it " + f"has to have a {' and a '.join(missing)} in it.", + ) + + layout = given["layout"] + if layout not in builtin_filters(): + raise _refuse( + lines["layout"], + f"{layout} is not a layout in2lambda has. The layouts are the filters: " + f"{', '.join(builtin_filters())}.", + ) + + strip = given.get("strip") or [] + if not isinstance(strip, list): + raise _refuse( + lines["strip"], f"strip is a list of patterns, which {strip!r} is not." + ) + return Spec( + question=_selector(given["question"], lines["question"]), + layout=layout, + part=_optional(given, "part", lines), + solution=_optional(given, "solution", lines), + ignore=_optional(given, "ignore", lines), + strip=[_pattern(pattern, lines["strip"]) for pattern in strip], + ) + + +def _optional( + given: dict[str, Any], name: str, lines: dict[str, int] +) -> Optional[Selector]: + """One of the selectors a spec need not have.""" + return _selector(given[name], lines[name]) if name in given else None + + +def _stems(roles: list[Optional[str]]) -> list[Optional[str]]: + """What each question and part is called - ``q1``, ``q1.a`` - in document order.""" + stems: list[Optional[str]] = [None] * len(roles) + questions, parts = 0, 0 + for index, role in enumerate(roles): + if role == "question": + questions, parts = questions + 1, 0 + stems[index] = f"q{questions}" + elif role == "part" and questions: + stems[index] = f"q{questions}.{chr(ord('a') + parts)}" + parts += 1 + return stems + + +def _slots(roles: list[Optional[str]], stems: list[Optional[str]]) -> list[str]: + """What a separate section of solutions answers, one after another. + + Each question with parts is its parts; each question without is itself. That is the + order the solutions in a PartsSepSol document are written in. + """ + questions: list[tuple[str, list[str]]] = [] + for index, role in enumerate(roles): + if role == "question": + questions.append((str(stems[index]), [])) + elif role == "part" and questions and stems[index]: + questions[-1][1].append(str(stems[index])) + return [slot for stem, parts in questions for slot in (parts or [stem])] + + +def _keys(layout: str, roles: list[Optional[str]]) -> list[Optional[str]]: + """The field each block's text goes in, or None where the layout puts it in none.""" + stems = _stems(roles) + separate = iter(_slots(roles, stems)) + keys: list[Optional[str]] = [None] * len(roles) + question: Optional[str] = None + parts: list[str] = [] + answered = 0 + for index, role in enumerate(roles): + if role == "question": + question, parts, answered = stems[index], [], 0 + keys[index] = f"{question}.text" + elif role == "part" and stems[index]: + parts.append(str(stems[index])) + keys[index] = f"{stems[index]}.text" + elif role == "solution" and question: + # Which solution answers what, mirroring the filter the layout names. The + # four filters are the four layouts; a fifth would need its rule adding. + target: Optional[str] + match layout: + case "PartsOneSol": # One solution to the whole question. + target = question + case "PartSolPartSol": # Each part answered where it stands. + target = parts[-1] if parts else question + case "PartPartSolSol": # The parts, then their solutions in order. + target = parts[answered] if answered < len(parts) else question + case _: # PartsSepSol: every solution together, at the end. + target = next(separate, None) + answered += 1 + keys[index] = f"{target}.solution" if target else None + return keys + + +def fields( + spec: Spec, elements: list[tuple[Block, Any]], markdown: str +) -> tuple[list[Field], list[str]]: + """What a spec makes of a document: its fields, and the blocks it says to ignore. + + Args: + spec: The spec to run, as :func:`load` read it. + elements: Every block of the frozen markdown beside the element it is, as + :func:`in2lambda.source._elements` gives them. + markdown: The frozen markdown itself, which the values are quoted out of. + + Returns: + One :class:`Field` per question, part and solution the spec found, and the ids + of the blocks its ``ignore`` selector matched. A block that is neither is in + neither, which is what a coverage report is about. + """ + import panflute as pf + + found = [element for _, element in elements] + roles = [ + next( + ( + role + for role in _ROLES + if (selector := getattr(spec, role)) is not None + and selector.matches(found, index, pf) + ), + None, + ) + for index in range(len(found)) + ] + lines = markdown.splitlines() + return ( + [ + Field(key, _stripped(spec, lines, block), [[block.start, block.end]]) + for (block, _), key in zip(elements, _keys(spec.layout, roles)) + if key is not None + ], + [block.id for (block, _), role in zip(elements, roles) if role == "ignore"], + ) + + +def _stripped(spec: Spec, lines: list[str], block: Block) -> str: + """A block's own lines of the source, with the spec's strip patterns taken off. + + The markdown rather than the text pandoc stringifies it to, so that the maths, the + emphasis and the images in a question survive into the field. + """ + text = "\n".join(lines[block.start - 1 : block.end]) + for pattern in spec.strip: + text = pattern.sub("", text) + return text.strip() diff --git a/poetry.lock b/poetry.lock index 0ce9238..3344e6d 100644 --- a/poetry.lock +++ b/poetry.lock @@ -1952,9 +1952,9 @@ files = [ packaging = ">=24.0" [extras] -convert = ["panflute"] +convert = ["panflute", "pyyaml"] [metadata] lock-version = "2.1" python-versions = "^3.10" -content-hash = "c66ed1332073cc43e889ad3646544b370b7bad03c41beb4952daf4a2623f62f5" +content-hash = "855ed560d278e045b39a71ede79f1c9058fab08b59e272a65abec9e4e34c47d8" diff --git a/pyproject.toml b/pyproject.toml index 3affe96..1dc30e3 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -24,6 +24,8 @@ documentation = "https://lambda-feedback.github.io/in2lambda" [tool.poetry.dependencies] python = "^3.10" panflute = { version = "^2.3.1", optional = true } +# Only for reading a spec, which is only ever run against a converted document. +pyyaml = { version = "^6.0", optional = true } rich-click = "^1.7.4" # At 0.20.0 and below, the import hook in in2lambda/__init__.py leaves `cli` a plain # function: `@cli.command()` then raises AttributeError at import (0.19.0, 0.20.0), or @@ -33,7 +35,7 @@ beartype = "^0.22" [tool.poetry.extras] # Only needed to convert documents; the Python API works without it. -convert = ["panflute"] +convert = ["panflute", "pyyaml"] [tool.poetry.scripts] in2lambda = "in2lambda.main:cli" diff --git a/tests/conftest.py b/tests/conftest.py index bbd5546..2812361 100644 --- a/tests/conftest.py +++ b/tests/conftest.py @@ -34,6 +34,12 @@ DRAFTS = sorted(path for path in DRAFTS_DIR.iterdir() if path.is_dir()) """Every draft folder, so that covering another command is a folder and no code.""" +SPECS_DIR = Path(__file__).parent / "fixtures" / "specs" +"""One document per folder, beside the spec to run over it and what it should make.""" + +SPECS = sorted(path for path in SPECS_DIR.iterdir() if path.is_dir()) +"""Every spec folder, so that covering another kind of document is a folder and no code.""" + @pytest.fixture def without_node(monkeypatch: pytest.MonkeyPatch) -> Iterator[None]: diff --git a/tests/fixtures/specs/README.md b/tests/fixtures/specs/README.md new file mode 100644 index 0000000..e3967e2 --- /dev/null +++ b/tests/fixtures/specs/README.md @@ -0,0 +1,23 @@ +# Specs run over a source + +Each folder here is one run of `in2lambda spec run`: a `source.md` to freeze, the `spec.yaml` to +run over it, the `expected.json` the spec should leave in the draft's `fields`, and the +`uncovered.txt` of the blocks the command should report as being in no field. The test freezes +the source, runs the spec, compares both, and then replays the draft from its log and checks the +file is unchanged byte for byte - so a folder covers what a spec makes of a document and that it +can be rebuilt from what was recorded. + +To cover another kind of document, add a folder. There is one per layout, since a layout is +only a rule about which solution answers which question or part: `parts_one_sol` has one +solution to each question, `part_sol_part_sol` a solution after each part, `part_part_sol_sol` +the parts and then their solutions in order, and `parts_sep_sol` every solution together at the +end. `parts_sep_sol` is the spec from the ticket, and leaves its `## Solutions` heading in no +field, which is what a coverage report is for. + +Each document is one the other three layouts read differently - a question with two parts and +one solution, or with a part nobody answered - since a document every layout agrees about +would pin no rule, and a layout given the wrong rule would go on passing. + +A selector matches what pandoc parses, and a field holds the markdown of the lines it was taken +from. That is why a part written `(a) Find the load.` is a `ListItem` - pandoc reads `(a)` as a +list marker - while `strip` still has to take the `(a) ` off the front of the value. diff --git a/tests/fixtures/specs/part_part_sol_sol/expected.json b/tests/fixtures/specs/part_part_sol_sol/expected.json new file mode 100644 index 0000000..ead90f0 --- /dev/null +++ b/tests/fixtures/specs/part_part_sol_sol/expected.json @@ -0,0 +1,110 @@ +{ + "b1.ignore": { + "by": "tests", + "edited": false, + "layer": 1, + "ranges": [ + [ + 1, + 1 + ] + ], + "value": true + }, + "q1.a.solution": { + "by": "tests", + "edited": false, + "layer": 1, + "ranges": [ + [ + 11, + 11 + ] + ], + "value": "$W = nRT\\ln(V_1/V_2)$." + }, + "q1.a.text": { + "by": "tests", + "edited": false, + "layer": 1, + "ranges": [ + [ + 5, + 5 + ] + ], + "value": "Find the work done on the gas." + }, + "q1.b.solution": { + "by": "tests", + "edited": false, + "layer": 1, + "ranges": [ + [ + 13, + 13 + ] + ], + "value": "$Q = W$, since the internal energy does not change." + }, + "q1.b.text": { + "by": "tests", + "edited": false, + "layer": 1, + "ranges": [ + [ + 7, + 7 + ] + ], + "value": "Find the heat rejected." + }, + "q1.c.text": { + "by": "tests", + "edited": false, + "layer": 1, + "ranges": [ + [ + 9, + 9 + ] + ], + "value": "Sketch the process on a $p$-$V$ diagram." + }, + "q1.text": { + "by": "tests", + "edited": false, + "layer": 1, + "ranges": [ + [ + 3, + 3 + ] + ], + "value": "An ideal gas is compressed isothermally from $V_1$ to $V_2$." + }, + "q2.solution": { + "by": "tests", + "edited": false, + "layer": 1, + "ranges": [ + [ + 17, + 17 + ] + ], + "value": "$\\eta = 1 - T_c/T_h = 0.5$." + }, + "q2.text": { + "by": "tests", + "edited": false, + "layer": 1, + "ranges": [ + [ + 15, + 15 + ] + ], + "value": "Find the efficiency of a Carnot engine between 300 K and 600 K." + } +} diff --git a/tests/fixtures/specs/part_part_sol_sol/source.md b/tests/fixtures/specs/part_part_sol_sol/source.md new file mode 100644 index 0000000..9f10886 --- /dev/null +++ b/tests/fixtures/specs/part_part_sol_sol/source.md @@ -0,0 +1,17 @@ +# Thermodynamics problem sheet + +Q1. An ideal gas is compressed isothermally from $V_1$ to $V_2$. + +(a) Find the work done on the gas. + +(b) Find the heat rejected. + +(c) Sketch the process on a $p$-$V$ diagram. + +Solution: $W = nRT\ln(V_1/V_2)$. + +Solution: $Q = W$, since the internal energy does not change. + +Q2. Find the efficiency of a Carnot engine between 300 K and 600 K. + +Solution: $\eta = 1 - T_c/T_h = 0.5$. diff --git a/tests/fixtures/specs/part_part_sol_sol/spec.yaml b/tests/fixtures/specs/part_part_sol_sol/spec.yaml new file mode 100644 index 0000000..8f455de --- /dev/null +++ b/tests/fixtures/specs/part_part_sol_sol/spec.yaml @@ -0,0 +1,6 @@ +question: Para text~'^Q\d+\.' +part: ListItem +solution: Para text~'^Solution:' +strip: ['^Q\d+\. ', '^\([a-z]\) ', '^Solution: '] +ignore: Header +layout: PartPartSolSol diff --git a/tests/fixtures/specs/part_part_sol_sol/uncovered.txt b/tests/fixtures/specs/part_part_sol_sol/uncovered.txt new file mode 100644 index 0000000..e69de29 diff --git a/tests/fixtures/specs/part_sol_part_sol/expected.json b/tests/fixtures/specs/part_sol_part_sol/expected.json new file mode 100644 index 0000000..0a29440 --- /dev/null +++ b/tests/fixtures/specs/part_sol_part_sol/expected.json @@ -0,0 +1,86 @@ +{ + "b1.ignore": { + "by": "tests", + "edited": false, + "layer": 1, + "ranges": [ + [ + 1, + 1 + ] + ], + "value": true + }, + "q1.a.text": { + "by": "tests", + "edited": false, + "layer": 1, + "ranges": [ + [ + 5, + 5 + ] + ], + "value": "Find the reaction at A." + }, + "q1.b.solution": { + "by": "tests", + "edited": false, + "layer": 1, + "ranges": [ + [ + 9, + 9 + ] + ], + "value": "$R_B = W/2$, and $R_A$ is the same by symmetry." + }, + "q1.b.text": { + "by": "tests", + "edited": false, + "layer": 1, + "ranges": [ + [ + 7, + 7 + ] + ], + "value": "Find the reaction at B." + }, + "q1.text": { + "by": "tests", + "edited": false, + "layer": 1, + "ranges": [ + [ + 3, + 3 + ] + ], + "value": "A uniform beam of weight $W$ rests on supports at A and B." + }, + "q2.solution": { + "by": "tests", + "edited": false, + "layer": 1, + "ranges": [ + [ + 13, + 13 + ] + ], + "value": "the sag is $wL^2/8T$." + }, + "q2.text": { + "by": "tests", + "edited": false, + "layer": 1, + "ranges": [ + [ + 11, + 11 + ] + ], + "value": "A cable of weight $w$ per metre hangs between two towers $L$ apart." + } +} diff --git a/tests/fixtures/specs/part_sol_part_sol/source.md b/tests/fixtures/specs/part_sol_part_sol/source.md new file mode 100644 index 0000000..e822c85 --- /dev/null +++ b/tests/fixtures/specs/part_sol_part_sol/source.md @@ -0,0 +1,13 @@ +# Statics problem sheet + +Q1. A uniform beam of weight $W$ rests on supports at A and B. + +(a) Find the reaction at A. + +(b) Find the reaction at B. + +Solution: $R_B = W/2$, and $R_A$ is the same by symmetry. + +Q2. A cable of weight $w$ per metre hangs between two towers $L$ apart. + +Solution: the sag is $wL^2/8T$. diff --git a/tests/fixtures/specs/part_sol_part_sol/spec.yaml b/tests/fixtures/specs/part_sol_part_sol/spec.yaml new file mode 100644 index 0000000..61cb788 --- /dev/null +++ b/tests/fixtures/specs/part_sol_part_sol/spec.yaml @@ -0,0 +1,6 @@ +question: Para text~'^Q\d+\.' +part: ListItem +solution: Para text~'^Solution:' +strip: ['^Q\d+\. ', '^\([a-z]\) ', '^Solution: '] +ignore: Header +layout: PartSolPartSol diff --git a/tests/fixtures/specs/part_sol_part_sol/uncovered.txt b/tests/fixtures/specs/part_sol_part_sol/uncovered.txt new file mode 100644 index 0000000..e69de29 diff --git a/tests/fixtures/specs/parts_one_sol/expected.json b/tests/fixtures/specs/parts_one_sol/expected.json new file mode 100644 index 0000000..bcd45ff --- /dev/null +++ b/tests/fixtures/specs/parts_one_sol/expected.json @@ -0,0 +1,110 @@ +{ + "b1.ignore": { + "by": "tests", + "edited": false, + "layer": 1, + "ranges": [ + [ + 1, + 1 + ] + ], + "value": true + }, + "b2.ignore": { + "by": "tests", + "edited": false, + "layer": 1, + "ranges": [ + [ + 3, + 3 + ] + ], + "value": true + }, + "b7.ignore": { + "by": "tests", + "edited": false, + "layer": 1, + "ranges": [ + [ + 13, + 13 + ] + ], + "value": true + }, + "q1.a.text": { + "by": "tests", + "edited": false, + "layer": 1, + "ranges": [ + [ + 7, + 7 + ] + ], + "value": "Find the thrust." + }, + "q1.b.text": { + "by": "tests", + "edited": false, + "layer": 1, + "ranges": [ + [ + 9, + 9 + ] + ], + "value": "Find the acceleration at launch." + }, + "q1.solution": { + "by": "tests", + "edited": false, + "layer": 1, + "ranges": [ + [ + 11, + 11 + ] + ], + "value": "the thrust is $\\dot m v_e$, so $a = \\dot m v_e / m - g$." + }, + "q1.text": { + "by": "tests", + "edited": false, + "layer": 1, + "ranges": [ + [ + 5, + 5 + ] + ], + "value": "A rocket burns fuel at a rate $\\dot m$ and exhausts it at $v_e$." + }, + "q2.solution": { + "by": "tests", + "edited": false, + "layer": 1, + "ranges": [ + [ + 17, + 17 + ] + ], + "value": "$c = \\sqrt{\\gamma R T}$." + }, + "q2.text": { + "by": "tests", + "edited": false, + "layer": 1, + "ranges": [ + [ + 15, + 15 + ] + ], + "value": "Find the speed of sound in air at 300 K." + } +} diff --git a/tests/fixtures/specs/parts_one_sol/source.md b/tests/fixtures/specs/parts_one_sol/source.md new file mode 100644 index 0000000..f8202a6 --- /dev/null +++ b/tests/fixtures/specs/parts_one_sol/source.md @@ -0,0 +1,17 @@ +# Momentum problem sheet + +## Rockets + +Q1. A rocket burns fuel at a rate $\dot m$ and exhausts it at $v_e$. + +(a) Find the thrust. + +(b) Find the acceleration at launch. + +Solution: the thrust is $\dot m v_e$, so $a = \dot m v_e / m - g$. + +## Sound + +Q2. Find the speed of sound in air at 300 K. + +Solution: $c = \sqrt{\gamma R T}$. diff --git a/tests/fixtures/specs/parts_one_sol/spec.yaml b/tests/fixtures/specs/parts_one_sol/spec.yaml new file mode 100644 index 0000000..fefb3cb --- /dev/null +++ b/tests/fixtures/specs/parts_one_sol/spec.yaml @@ -0,0 +1,6 @@ +question: Para text~'^Q\d+\.' +part: ListItem +solution: Para text~'^Solution:' +strip: ['^Q\d+\. ', '^\([a-z]\) ', '^Solution: '] +ignore: Header +layout: PartsOneSol diff --git a/tests/fixtures/specs/parts_one_sol/uncovered.txt b/tests/fixtures/specs/parts_one_sol/uncovered.txt new file mode 100644 index 0000000..e69de29 diff --git a/tests/fixtures/specs/parts_sep_sol/expected.json b/tests/fixtures/specs/parts_sep_sol/expected.json new file mode 100644 index 0000000..02e4286 --- /dev/null +++ b/tests/fixtures/specs/parts_sep_sol/expected.json @@ -0,0 +1,98 @@ +{ + "b1.ignore": { + "by": "tests", + "edited": false, + "layer": 1, + "ranges": [ + [ + 1, + 1 + ] + ], + "value": true + }, + "q1.a.solution": { + "by": "tests", + "edited": false, + "layer": 1, + "ranges": [ + [ + 13, + 13 + ] + ], + "value": "The load is $F = pA$." + }, + "q1.a.text": { + "by": "tests", + "edited": false, + "layer": 1, + "ranges": [ + [ + 5, + 5 + ] + ], + "value": "Find the load the large piston carries." + }, + "q1.b.solution": { + "by": "tests", + "edited": false, + "layer": 1, + "ranges": [ + [ + 15, + 15 + ] + ], + "value": "The pressure is $p = F/a$." + }, + "q1.b.text": { + "by": "tests", + "edited": false, + "layer": 1, + "ranges": [ + [ + 7, + 7 + ] + ], + "value": "Find the pressure in the oil." + }, + "q1.text": { + "by": "tests", + "edited": false, + "layer": 1, + "ranges": [ + [ + 3, + 3 + ] + ], + "value": "Question 1: a hydraulic scale with two pistons joined by oil" + }, + "q2.solution": { + "by": "tests", + "edited": false, + "layer": 1, + "ranges": [ + [ + 17, + 17 + ] + ], + "value": "The drag is $\\tau A$." + }, + "q2.text": { + "by": "tests", + "edited": false, + "layer": 1, + "ranges": [ + [ + 9, + 9 + ] + ], + "value": "Question 2: a flat plate towed through water at $u$" + } +} diff --git a/tests/fixtures/specs/parts_sep_sol/source.md b/tests/fixtures/specs/parts_sep_sol/source.md new file mode 100644 index 0000000..bccb6d1 --- /dev/null +++ b/tests/fixtures/specs/parts_sep_sol/source.md @@ -0,0 +1,17 @@ +# Fluid mechanics problem sheet + +## Question 1: a hydraulic scale with two pistons joined by oil + +(a) Find the load the large piston carries. + +(b) Find the pressure in the oil. + +## Question 2: a flat plate towed through water at $u$ + +## Solutions + +1(a) The load is $F = pA$. + +1(b) The pressure is $p = F/a$. + +2 The drag is $\tau A$. diff --git a/tests/fixtures/specs/parts_sep_sol/spec.yaml b/tests/fixtures/specs/parts_sep_sol/spec.yaml new file mode 100644 index 0000000..2c058ae --- /dev/null +++ b/tests/fixtures/specs/parts_sep_sol/spec.yaml @@ -0,0 +1,6 @@ +question: Header level=2 text~'^Question' +part: ListItem +solution: after Header text=Solutions, label~'^\d+(\([a-z]\))?$' +strip: ['^#+ ', '^\([a-z]\) ', '^\d+(\([a-z]\))? '] +ignore: Header level=1 +layout: PartsSepSol diff --git a/tests/fixtures/specs/parts_sep_sol/uncovered.txt b/tests/fixtures/specs/parts_sep_sol/uncovered.txt new file mode 100644 index 0000000..07eb61d --- /dev/null +++ b/tests/fixtures/specs/parts_sep_sol/uncovered.txt @@ -0,0 +1 @@ +b6 diff --git a/tests/test_cli.py b/tests/test_cli.py index 56515d0..491cb2e 100644 --- a/tests/test_cli.py +++ b/tests/test_cli.py @@ -89,4 +89,5 @@ def test_completing_the_old_form_offers_the_subcommand() -> None: "convert", "draft", "source", + "spec", ] diff --git a/tests/test_spec.py b/tests/test_spec.py new file mode 100644 index 0000000..cc3699f --- /dev/null +++ b/tests/test_spec.py @@ -0,0 +1,298 @@ +"""Running a spec of selectors over a frozen source. + +Each folder in ``fixtures/specs`` is a document, the spec to run over it, the fields it +should fill in and the blocks it should leave out, so covering another kind of document +means adding a folder rather than a test. The rest is what the command line does with a +spec it cannot read - a typo in a selector, a layout nothing has, a file edited since it +was run - which is not something a fixture can say. +""" + +import hashlib +import json +import shutil +import sys +from pathlib import Path +from typing import Any + +import pytest +from click.testing import CliRunner +from conftest import SPECS, SPECS_DIR + +from in2lambda.main import cli + +WORKED_EXAMPLE = SPECS_DIR / "parts_sep_sol" +"""The case the tests below happen to use; what they check holds for any of them.""" + + +def _frozen(folder: Path, tmp_path: Path) -> CliRunner: + """A folder's document and its spec, copied into `tmp_path` with the source frozen.""" + shutil.copytree(folder, tmp_path, dirs_exist_ok=True) + runner = CliRunner() + assert runner.invoke(cli, ["source", "add", "source.md"]).exit_code == 0 + return runner + + +@pytest.mark.parametrize("folder", SPECS, ids=lambda path: path.name) +def test_a_spec_fills_in_the_fields_beside_it_and_replays( + folder: Path, tmp_path: Path, monkeypatch +) -> None: + """What a spec makes of a document, and that its log rebuilds the same draft.""" + monkeypatch.chdir(tmp_path) + runner = _frozen(folder, tmp_path) + + result = runner.invoke(cli, ["spec", "run", "spec.yaml", "--by", "tests"]) + + assert result.exit_code == 0, result.output + draft_path = tmp_path / "draft.json" + draft = json.loads(draft_path.read_text()) + assert draft["fields"] == json.loads((folder / "expected.json").read_text()) + + reported = [ + line.split()[0] for line in result.output.splitlines() if "no field" in line + ] + assert reported == (folder / "uncovered.txt").read_text().split() + + # The spec is named and hashed in the log, so a replay runs the one that ran. + spec = (tmp_path / "spec.yaml").read_bytes() + assert draft["log"] == [ + { + "command": "spec run", + "args": { + "spec": "spec.yaml", + "hash": f"sha256:{hashlib.sha256(spec).hexdigest()}", + }, + "by": "tests", + } + ] + + written = draft_path.read_bytes() + replay = runner.invoke(cli, ["draft", "replay"]) + + assert replay.exit_code == 0, replay.output + assert draft_path.read_bytes() == written + + +def test_a_spec_ignoring_a_block_the_draft_has_split_covers_both_halves( + tmp_path: Path, monkeypatch +) -> None: + """`split block` gives the draft blocks the parser, which a spec runs over, has not.""" + monkeypatch.setenv("COLUMNS", "200") + monkeypatch.chdir(tmp_path) + (tmp_path / "source.md").write_text( + "Instructions: answer every question.\nThey are not marked.\n\n" + "Q1. Find the load the large piston carries.\n\nSolution: $F = pA$.\n" + ) + (tmp_path / "spec.yaml").write_text( + "question: Para text~'^Q\\d+\\.'\n" + "solution: Para text~'^Solution:'\n" + "strip: ['^Q\\d+\\. ', '^Solution: ']\n" + "ignore: Para text~'^Instructions'\n" + "layout: PartsOneSol\n" + ) + runner = CliRunner() + assert runner.invoke(cli, ["source", "add", "source.md"]).exit_code == 0 + # The two sentences are one paragraph to pandoc, so the draft now has b1a and b1b + # where the spec, run over the source again, sees the one block b1. + assert runner.invoke(cli, ["draft", "split", "block", "b1", "2"]).exit_code == 0 + + result = runner.invoke(cli, ["spec", "run", "spec.yaml", "--by", "tests"]) + + assert result.exit_code == 0, result.output + fields = json.loads((tmp_path / "draft.json").read_text())["fields"] + assert fields["b1.ignore"]["ranges"] == [[1, 2]] + # Both halves are within the lines the ignore field was written over, so neither is + # reported as left out. + assert "is in no field" not in result.output + + +def test_a_replay_is_refused_once_the_spec_has_changed( + tmp_path: Path, monkeypatch +) -> None: + """The fields came from the spec as it was, so a replay of a new one proves nothing.""" + monkeypatch.setenv("COLUMNS", "200") # So the message is not wrapped mid-sentence. + monkeypatch.chdir(tmp_path) + runner = _frozen(WORKED_EXAMPLE, tmp_path) + assert runner.invoke(cli, ["spec", "run", "spec.yaml"]).exit_code == 0 + draft_path = tmp_path / "draft.json" + written = draft_path.read_bytes() + + spec = tmp_path / "spec.yaml" + spec.write_text(spec.read_text().replace("PartsSepSol", "PartsOneSol")) + result = runner.invoke(cli, ["draft", "replay"]) + + assert result.exit_code != 0 + assert "spec.yaml has changed" in result.output + assert draft_path.read_bytes() == written + + +@pytest.mark.parametrize( + "named", + ["spec.yaml", "./spec.yaml", "spec2.yaml"], + ids=["as it was", "spelled another way", "as a copy"], +) +def test_running_an_edited_spec_again_is_refused( + named: str, tmp_path: Path, monkeypatch +) -> None: + """The fields of the first run would stay, and the draft could never replay again.""" + monkeypatch.setenv("COLUMNS", "200") + monkeypatch.chdir(tmp_path) + runner = _frozen(WORKED_EXAMPLE, tmp_path) + assert runner.invoke(cli, ["spec", "run", "spec.yaml"]).exit_code == 0 + draft_path = tmp_path / "draft.json" + written = draft_path.read_bytes() + + spec = tmp_path / "spec.yaml" + spec.write_text(spec.read_text().replace("PartsSepSol", "PartsOneSol")) + # However the second run names the spec - the way the first did, another way round + # to the same file, or as a copy under a name of its own - what is refused is that + # the spec the draft was filled in from has changed, since that is what no replay + # could get past afterwards. + shutil.copy(spec, tmp_path / "spec2.yaml") + result = runner.invoke(cli, ["spec", "run", named]) + + assert result.exit_code != 0 + # Named as the spec that ran, whatever this run called it, and refused for having + # changed rather than for the fields of the first run being in the way. + assert "spec.yaml has changed" in result.output + assert "--start-over" in result.output + assert draft_path.read_bytes() == written + # And what is on disk is still a draft that replays, which is the point of refusing. + spec.write_text(spec.read_text().replace("PartsOneSol", "PartsSepSol")) + assert runner.invoke(cli, ["draft", "replay"]).exit_code == 0 + + +@pytest.mark.parametrize("entry", [5, "nonsense"], ids=["a number", "some words"]) +def test_a_spec_run_over_a_log_holding_something_that_is_not_a_command_is_refused( + entry: Any, tmp_path: Path, monkeypatch +) -> None: + """A spec run reads the log it adds to, which is a file anyone can have edited.""" + monkeypatch.setenv("COLUMNS", "200") + monkeypatch.chdir(tmp_path) + runner = _frozen(WORKED_EXAMPLE, tmp_path) + assert runner.invoke(cli, ["spec", "run", "spec.yaml"]).exit_code == 0 + draft_path = tmp_path / "draft.json" + draft = json.loads(draft_path.read_text()) + draft["log"].append(entry) + draft_path.write_text(json.dumps(draft)) + written = draft_path.read_bytes() + + result = runner.invoke(cli, ["spec", "run", "spec.yaml"]) + + assert result.exit_code != 0 + # The same thing `draft replay` says of the same log, rather than a traceback from + # whichever line indexed it first. + assert "is not a command" in result.output + assert isinstance(result.exception, SystemExit) + assert draft_path.read_bytes() == written + + +def test_a_replay_is_refused_once_the_spec_has_gone( + tmp_path: Path, monkeypatch +) -> None: + """A draft names the spec that filled it in, and someone may well have moved it.""" + monkeypatch.setenv("COLUMNS", "200") + monkeypatch.chdir(tmp_path) + runner = _frozen(WORKED_EXAMPLE, tmp_path) + assert runner.invoke(cli, ["spec", "run", "spec.yaml"]).exit_code == 0 + (tmp_path / "spec.yaml").unlink() + + result = runner.invoke(cli, ["draft", "replay"]) + + assert result.exit_code != 0 + assert "spec.yaml" in result.output + + +@pytest.mark.parametrize( + ("spec", "line", "named"), + [ + ("question: Header\n layout: PartsOneSol\n", "line 2", "not YAML"), + ("question: !Header\nlayout: PartsOneSol\n", "line 1", "not YAML"), + ("question: Header\n? [a, b]\n: Header\n", "line 2", "unhashable"), + ("question: Header\nlayout: Sausage\n", "line 2", "Sausage"), + ("question: Sausage\nlayout: PartsOneSol\n", "line 1", "pandoc element"), + ("question: Header colour=blue\nlayout: PartsOneSol\n", "line 1", "colour"), + ("quesiton: Header\nlayout: PartsOneSol\n", "line 1", "quesiton"), + ], + ids=[ + "not yaml", + "tag nothing constructs", + "key nothing can hash", + "unknown layout", + "unknown type", + "unknown attribute", + "typo", + ], +) +def test_a_spec_that_cannot_be_read_says_which_line_to_look_at( + spec: str, line: str, named: str, tmp_path: Path, monkeypatch +) -> None: + """A spec is written by hand, so a mistake in one is a message, not a traceback.""" + monkeypatch.setenv("COLUMNS", "200") + monkeypatch.chdir(tmp_path) + runner = _frozen(WORKED_EXAMPLE, tmp_path) + written = (tmp_path / "draft.json").read_bytes() + (tmp_path / "spec.yaml").write_text(spec) + + result = runner.invoke(cli, ["spec", "run", "spec.yaml"]) + + assert result.exit_code != 0 + assert named in result.output + assert line in result.output + assert isinstance(result.exception, SystemExit) + # Nothing is half written: the draft is as it was before the spec was run. + assert (tmp_path / "draft.json").read_bytes() == written + + +def test_a_spec_saved_as_utf_16_is_read_like_any_other( + tmp_path: Path, monkeypatch +) -> None: + """A spec is written in whatever the editor saves in, and YAML reads the BOM.""" + monkeypatch.chdir(tmp_path) + runner = _frozen(WORKED_EXAMPLE, tmp_path) + spec = tmp_path / "spec.yaml" + spec.write_bytes(spec.read_text().encode("utf-16")) + + result = runner.invoke(cli, ["spec", "run", "spec.yaml", "--by", "tests"]) + + assert result.exit_code == 0, result.output + fields = json.loads((tmp_path / "draft.json").read_text())["fields"] + assert fields == json.loads((WORKED_EXAMPLE / "expected.json").read_text()) + + +def test_a_spec_in_an_encoding_yaml_cannot_read_is_refused( + tmp_path: Path, monkeypatch +) -> None: + """A spec saved as cp1252 is something to say so about, not a decoding traceback.""" + monkeypatch.setenv("COLUMNS", "200") + monkeypatch.chdir(tmp_path) + runner = _frozen(WORKED_EXAMPLE, tmp_path) + written = (tmp_path / "draft.json").read_bytes() + (tmp_path / "spec.yaml").write_bytes( + "question: Header\nstrip: ['^Solución ']\nlayout: PartsOneSol\n".encode( + "cp1252" + ) + ) + + result = runner.invoke(cli, ["spec", "run", "spec.yaml"]) + + assert result.exit_code != 0 + assert "not YAML" in result.output + assert isinstance(result.exception, SystemExit) + assert (tmp_path / "draft.json").read_bytes() == written + + +def test_running_a_spec_without_pyyaml_says_what_to_install( + tmp_path: Path, monkeypatch +) -> None: + """Reading a spec needs pyyaml, which only the convert extra installs.""" + monkeypatch.setenv("COLUMNS", "200") + monkeypatch.chdir(tmp_path) + runner = _frozen(WORKED_EXAMPLE, tmp_path) + monkeypatch.setitem(sys.modules, "yaml", None) + + result = runner.invoke(cli, ["spec", "run", "spec.yaml"]) + + assert result.exit_code != 0 + assert "pyyaml" in result.output + assert "pip install 'in2lambda[convert]'" in result.output + assert isinstance(result.exception, SystemExit)