From fe91a9742a537e905694a15a6e4e784b27645e00 Mon Sep 17 00:00:00 2001 From: Claude Date: Mon, 28 Sep 2026 19:57:00 +0000 Subject: [PATCH] Filesystem probing, capability profiling and on-demand reproduction Rebased onto master's known-issues rearchitecture (#14). The previous 21 commits were replayed as one: they were written against the old layout, and replaying them individually meant re-resolving the same files against text that no longer exists (per-cell dispatcher workflows, `.github/matrix.yaml`, a KNOWN_RED dict). The pre-rebase tip is kept locally at refs/backup/pre-rebase-pr5 (f704576). What this adds: - bin/ci/probe-backend.sh, .github/workflows/probe-filesystems.yaml and bin/ci/render-probe-table.sh: try to stand a candidate filesystem up on a runner and report what it got, over 24 candidates. "NOT BOOTSTRAPPABLE" is a finding, so a verdict exits 0 while a genuine script error does not. - bin/ci/fs-capabilities.sh: what a mounted filesystem supports in the dimensions git-annex trips over, as key=value lines. Ten seconds, and it usually explains a suite result before the suite is worth starting: hardlink-same-inode=no predicts `failed to link to annex`. - The `capabilities` target (bin/ci/target-capabilities.sh) plus .github/workflows/reproduce.yaml: the support entry point, for running one combination now with a reporter's own mount options. - FILESYSTEMS.md: reported breakage, then what the probes measured here, then feasibility. - bin/eval-under-loop gains gfs2, ocfs2, f2fs, exfat and nilfs2, a --mkfs-opts override, per-filesystem size floors, and integer validation of --size (a non-numeric value previously reached `dd count=40M`, i.e. 40 TiB). Ported to the new architecture rather than force-fitted: - The matrix lives in evals/matrix.yaml now, so the `capabilities` target and its `on-demand` flag go there. - `on-demand` filtering moved from copies in update-status.py and render-report.py into the one place cells are now enumerated, matrix_cells() in bin/ci/evals.py, with render-report.py dropping the column. Verified: 20 scheduled cells, no capabilities cell, no capabilities column. - GOTCHAS.md's image-size row now describes `loop-size-mb` in evals/matrix.yaml rather than the retired target_loop_size_mb(). - README's file-layout table is master's, plus this branch's six rows; the old table it carried still listed the deleted per-cell dispatchers. Verified: bin/ci/run-checks.sh passes in full -- shellcheck over 29 scripts, pyflakes, bats (28), known-issues validate, 28 unit tests. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01E1fDiGVWKywpbhVps5hzQK --- .github/workflows/probe-filesystems.yaml | 147 +++++++ .github/workflows/reproduce.yaml | 127 ++++++ FILESYSTEMS.md | 371 ++++++++++++++++ GOTCHAS.md | 21 +- README.md | 45 ++ bin/ci/evals.py | 10 +- bin/ci/fs-capabilities.sh | 234 ++++++++++ bin/ci/install-backend.sh | 19 + bin/ci/install-target.sh | 6 +- bin/ci/matrix.sh | 18 +- bin/ci/probe-backend.sh | 529 +++++++++++++++++++++++ bin/ci/render-probe-table.sh | 53 +++ bin/ci/render-report.py | 4 +- bin/ci/run-under.sh | 13 +- bin/ci/target-capabilities.sh | 28 ++ bin/eval-under-loop | 101 ++++- evals/matrix.yaml | 18 + 17 files changed, 1730 insertions(+), 14 deletions(-) create mode 100644 .github/workflows/probe-filesystems.yaml create mode 100644 .github/workflows/reproduce.yaml create mode 100644 FILESYSTEMS.md create mode 100755 bin/ci/fs-capabilities.sh create mode 100755 bin/ci/probe-backend.sh create mode 100755 bin/ci/render-probe-table.sh create mode 100755 bin/ci/target-capabilities.sh diff --git a/.github/workflows/probe-filesystems.yaml b/.github/workflows/probe-filesystems.yaml new file mode 100644 index 0000000..42ec5dc --- /dev/null +++ b/.github/workflows/probe-filesystems.yaml @@ -0,0 +1,147 @@ +name: Probe candidate filesystems + +# Reconnaissance, not a test matrix: for each candidate filesystem, try +# to stand up a throwaway instance on a stock GitHub-hosted runner and +# report what it supports. Feeds the survey in FILESYSTEMS.md and +# answers "is a bin/eval-under- backend even possible here?" +# before anyone writes one. +# +# Every candidate reports a verdict and the job stays green: "could not +# be brought up on this runner" is the answer to the question, not a +# failure. Read the roll-up table in the run summary -- the check mark +# only tells you the probe itself ran. +# +# All logic lives in bin/ci/probe-backend.sh -- see .claude/CLAUDE.md. + +# Reconnaissance is 28 jobs, so it does not run on every branch push. +# Dispatch it when you want an answer; master re-runs it when the probe +# itself changes, to catch bit-rot in the bring-up recipes. +on: + workflow_dispatch: + push: + branches: [master] + paths: + - "bin/ci/probe-backend.sh" + - "bin/ci/fs-capabilities.sh" + - ".github/workflows/probe-filesystems.yaml" + +concurrency: + group: probe-filesystems-${{ github.ref }} + cancel-in-progress: true + +permissions: + contents: read + +jobs: + probe: + name: ${{ matrix.candidate }} (${{ matrix.os }}) + runs-on: ${{ matrix.os }} + timeout-minutes: 25 + strategy: + fail-fast: false + matrix: + os: [ubuntu-24.04] + candidate: + - gfs2 + - ocfs2 + - f2fs + - exfat + - ntfs3 + - udf + - nilfs2 + - bcachefs + - zfs + - overlay + - glusterfs + - cephfs + - cifs + - sshfs + - gocryptfs + - encfs + - ecryptfs + - s3-rclone + - lustre-client + - openafs + - vm-only + include: + # Whamcloud publishes Lustre client packages per Ubuntu LTS, + # and 22.04 is the one with the longest support matrix -- so + # the client-module question gets asked on both images. The + # two cluster filesystems get the same treatment because their + # kernel modules live in linux-modules-extra, whose contents + # differ between runner images. + - os: ubuntu-22.04 + candidate: lustre-client + - os: ubuntu-22.04 + candidate: gfs2 + - os: ubuntu-22.04 + candidate: ocfs2 + + steps: + - uses: actions/checkout@v4 + + - name: Probe ${{ matrix.candidate }} + run: | + set -o pipefail + bin/ci/probe-backend.sh "${{ matrix.candidate }}" \ + 2>&1 | tee "$RUNNER_TEMP/probe.log" + + # Runs even when the probe failed: "could not be brought up, here + # is why" is the finding, and it belongs in the roll-up too. + - name: Extract the summary line + if: always() + run: | + grep '^SUMMARY|' "$RUNNER_TEMP/probe.log" > "$RUNNER_TEMP/summary.txt" \ + || echo "SUMMARY|${{ matrix.candidate }}|NO OUTPUT|-|probe produced no summary line|" \ + > "$RUNNER_TEMP/summary.txt" + sed -i "s/^SUMMARY|/SUMMARY|${{ matrix.os }}|/" "$RUNNER_TEMP/summary.txt" + cat "$RUNNER_TEMP/summary.txt" + + - name: Upload the summary line + if: always() + uses: actions/upload-artifact@v4 + with: + name: summary-${{ matrix.os }}-${{ matrix.candidate }} + path: ${{ runner.temp }}/summary.txt + retention-days: 30 + + # The probe answers "can this filesystem exist here?". This answers the + # follow-up: does the *backend* -- the thing a matrix row would call -- + # actually drive it end to end? ext4 is the control. + backend-smoke: + name: eval-under loop --fs ${{ matrix.fs }} + runs-on: ubuntu-24.04 + timeout-minutes: 15 + strategy: + fail-fast: false + matrix: + fs: [ext4, gfs2, ocfs2] + + steps: + - uses: actions/checkout@v4 + + - name: Install backend deps + run: bin/ci/install-backend.sh loop "${{ matrix.fs }}" + + - name: Run the capability probe through the backend + run: | + sudo bin/eval-under loop --fs "${{ matrix.fs }}" --set-home -- \ + bin/ci/fs-capabilities.sh + + # One place to read the whole run. Without this, answering "which + # candidates came up?" means opening 24 job logs. + roll-up: + name: Roll-up + runs-on: ubuntu-24.04 + needs: [probe, backend-smoke] + if: always() + steps: + - uses: actions/checkout@v4 + - uses: actions/download-artifact@v4 + with: + pattern: summary-* + path: summaries + merge-multiple: false + + - name: Render the table + run: bin/ci/render-probe-table.sh summaries | tee -a "$GITHUB_STEP_SUMMARY" diff --git a/.github/workflows/reproduce.yaml b/.github/workflows/reproduce.yaml new file mode 100644 index 0000000..d8177e0 --- /dev/null +++ b/.github/workflows/reproduce.yaml @@ -0,0 +1,127 @@ +name: Reproduce (on demand) + +# The support entry point: someone reports "X breaks on my $FILESYSTEM", +# and you want a run of exactly that combination now -- without waiting +# for the Monday schedule and without adding a matrix cell for a +# filesystem we are not going to watch every week. +# +# Deliberately outside evals/matrix.yaml: no badge, no cron, no push +# trigger. Every field is a dispatch input, including targets +# (capabilities) that have no scheduled cells at all. +# +# The job body mirrors test.yaml's, minus the matrix fan-out and the +# status publishing -- a reproduction must never rewrite the status that +# the README and the report page show for master. +# +# All non-trivial shell logic lives in bin/ci/*.sh. See .claude/CLAUDE.md. + +on: + workflow_dispatch: + inputs: + backend: + description: "Filesystem to run under" + required: true + type: choice + default: nfs + options: [nfs, loop, beegfs] + backend-version: + description: >- + loop: filesystem (ext4, vfat, xfs, btrfs, gfs2, ocfs2, f2fs, + exfat, nilfs2, bcachefs). beegfs: point release (7.4.6, 8.1.0). + nfs: leave as n/a. + required: true + type: string + default: "n/a" + target: + description: "What to run under it" + required: true + type: choice + default: capabilities + options: [capabilities, git-annex, git, stress-ng, pjdfstest] + backend-options: + description: >- + Extra backend flags, e.g. "--sync" or "--no-root-squash" + (nfs). Run `bin/eval-under BACKEND --help` + for the list. + required: false + type: string + default: "" + runner: + description: "Runner image (beegfs requires ubuntu-22.04)" + required: true + type: choice + default: ubuntu-24.04 + options: [ubuntu-24.04, ubuntu-22.04] + +# Named for the combination, so two people reproducing two different +# reports do not cancel each other. +concurrency: + group: >- + reproduce-${{ inputs.backend }}-${{ inputs.backend-version }}-${{ inputs.target }}-${{ github.ref }} + cancel-in-progress: false + +jobs: + reproduce: + name: ${{ inputs.target }} under ${{ inputs.backend }} ${{ inputs.backend-version }} + runs-on: ${{ inputs.runner }} + timeout-minutes: 60 + permissions: + contents: read + actions: read # to download artifacts from con/git-annex + steps: + - uses: actions/checkout@v4 + + # Every input reaches the shell through env, never through template + # substitution into the run: body. A dispatch input is free-form + # text, and ${{ }} is expanded before bash sees the line, so + # interpolating it directly would make the input executable. This + # also matches how run-under.sh already takes backend-options. + - name: Install backend client dependencies + env: + BACKEND: ${{ inputs.backend }} + BACKEND_VERSION: ${{ inputs.backend-version }} + run: bin/ci/install-backend.sh "$BACKEND" "$BACKEND_VERSION" + + - name: Install target ${{ inputs.target }} + env: + TARGET: ${{ inputs.target }} + run: bin/ci/install-target.sh "$TARGET" + + - name: Fetch latest git-annex daily build from con/git-annex + if: inputs.target == 'git-annex' + env: + GH_TOKEN: ${{ secrets.GITHUB_TOKEN }} + run: bin/ci/install-git-annex-daily.sh + + - name: Configure git identity + run: | + git config --global user.email test@github.land + git config --global user.name "GitHub Almighty" + + - name: Run ${{ inputs.target }} under ${{ inputs.backend }} + env: + BACKEND: ${{ inputs.backend }} + BACKEND_VERSION: ${{ inputs.backend-version }} + TARGET: ${{ inputs.target }} + EVAL_UNDER_BACKEND_OPTS: ${{ inputs.backend-options }} + run: >- + sudo -E bin/ci/run-under.sh "$BACKEND" "$BACKEND_VERSION" "$TARGET" + + - name: Dump failure logs + if: failure() + env: + BACKEND: ${{ inputs.backend }} + BACKEND_VERSION: ${{ inputs.backend-version }} + TARGET: ${{ inputs.target }} + run: >- + bin/ci/dump-failure-logs.sh "$BACKEND" "$BACKEND_VERSION" "$TARGET" + + - name: Upload logs + if: always() + uses: actions/upload-artifact@v4 + with: + name: reproduce-logs-${{ inputs.backend }}-${{ inputs.target }}-${{ github.run_id }} + path: | + /var/log/beegfs-* + /opt/eval-under-src/git/t/test-results/** + if-no-files-found: ignore diff --git a/FILESYSTEMS.md b/FILESYSTEMS.md new file mode 100644 index 0000000..c23f6bb --- /dev/null +++ b/FILESYSTEMS.md @@ -0,0 +1,371 @@ +# Which filesystem next? + +This repo can run any suite under any filesystem it can mount. That +raises two separate questions, and they have different answers: + +1. **Which filesystems have actually broken git-annex or DataLad for + real users?** -- answered from bug trackers, below. +2. **Which of those can we stand up in CI?** -- answered by *measuring* + it: `bin/ci/probe-backend.sh` tries to bring a candidate up on a stock + GitHub-hosted runner and reports what it got. + +A filesystem is worth a matrix row when both answers are yes. This file +records both columns so the next person does not redo either. + +> Re-run the measurements any time: the **Probe candidate filesystems** +> workflow (`workflow_dispatch`, or push a change to +> `bin/ci/probe-backend.sh`). Its roll-up job prints the whole table. + +## Part 1 -- what users actually hit + +Sorted by how much evidence there is, not by how exotic the filesystem +is. The "breaks" column names the *mechanism*, because the mechanism is +what a test has to reproduce. + +| Filesystem | What breaks | Evidence | +| --- | --- | --- | +| **NFS** | POSIX record locks unreliable or absent, so `git-annex` falls back to `annex.pidlock`; close-to-open consistency makes freshly-created files invisible for seconds; deleting an open file silly-renames to `.nfsXXXX` | [`waitToSetLock: resource exhausted (No locks available)`][nfs1], [infinite hang on `git status` with v10][nfs2], [`git annex drop` does not free space][nfs3], DataLad [xfails archive tests on NFS][nfs4] and [gives NFS runs 20x the time budget][nfs5] | +| **Lustre** | SQLite in WAL mode returns `disk I/O error`; POSIX locks missing, so pidlock again; and `link()` **succeeds over an existing name**, leaving two directory entries with one name -- which is what pidlock relies on being atomic | [drop blows on lustre: SQLite3 returned ErrorIO][lus1], [day 336: pid locks][lus2], [day 337: who needs POSIX][lus3] | +| **BeeGFS** | rename semantics; 35+ test failures on 7.4.6 | [35 failed tests on beegfs][bee1] -- the report this repo was built for | +| **CIFS / SMB** | SQLite locking fails against the server's ownership model (workaround: mount `nobrl`); `git annex init` cannot remove a named pipe with unix extensions on; case-insensitive | [git-annex on Samba share][smb1] | +| **vfat / exFAT / NTFS / WSL DrvFs** | no symlinks, no ownership, no exec bit -- git-annex switches to an adjusted branch, and that path has its own bugs | [WSL1: git-annex-add fails in DrvFs][crip1], [WSL adjusted branches: smudge fails, sqlite locking protocol][crip2], [crippled fs (pidlock) leads to SQLite3 error][crip3], [files unaccessible in views on a crippled filesystem][crip4], DataLad [#258][crip5], [#4777][crip6] | +| **Isilon (NFS server)** | `cp -a` preserving the server's xattrs breaks the suite; 357/984 tests failed | ["357 out of 984 tests failed"][iso1] | +| **GlusterFS** | SQLite WAL does not work on Gluster -- the same mechanism as Lustre, reported on the Gluster side | [Gluster-users: Locking and SQLite][glu1] | +| **GPFS / IBM Storage Scale** | inode budget exhaustion with many annexed files; transfer-lock errors reported on HPC | [DataLad #5589][gpf1], [handbook: HPC installation notes][gpf2]. *No first-party git-annex bug report found* -- the GPFS reports are user-forum-shaped, which is itself a finding: nobody has run the suite there. | +| **ZFS** | not a breakage: `cp --reflink` unsupported, and `copy_file_range` behaviour over NFS+ZFS is a known open design question | [use copy_file_range for get and copy][zfs1] | + +[nfs1]: https://git-annex.branchable.com/projects/datalad/bugs-done/git_annex_info_fails_on_NFS__58___waitToSetLock__58___resource_exhausted___40__No_locks_available__41__/ +[nfs2]: https://git-annex.branchable.com/bugs/infinite_hang_on_git_status_with_v10___38___nfs/ +[nfs3]: https://github.com/datalad/datalad/issues/3929 +[nfs4]: https://github.com/datalad/datalad/pull/6912 +[nfs5]: https://github.com/datalad/datalad/pull/7749 +[lus1]: https://git-annex.branchable.com/bugs/drop_blows_on_lustre__58___SQLite3_returned_ErrorIO/ +[lus2]: https://git-annex.branchable.com/devblog/day_336__pid_locks/ +[lus3]: https://git-annex.branchable.com/devblog/day_337__who_needs_POSIX/ +[bee1]: https://git-annex.branchable.com/bugs/35_failed_tests_on_beegfs/ +[smb1]: https://git-annex.branchable.com/forum/git-annex_on_Samba_share/ +[crip1]: https://git-annex.branchable.com/bugs/WSL1__58___git-annex-add_fails_in_DrvFs_filesystem/ +[crip2]: https://git-annex.branchable.com/bugs/WSL_adjusted_braches__58___smudge_fails_with_sqlite_thread_crashed_-_locking_protocol/ +[crip3]: https://git-annex.branchable.com/bugs/crippled_fs___40__pidlock__41___leads_to_git-annex__58___SQLite3_error/ +[crip4]: https://git-annex.branchable.com/bugs/Files_unaccessible_in___40__some__63____41___views_on_a_crippled_filesystem/ +[crip5]: https://github.com/datalad/datalad/issues/258 +[crip6]: https://github.com/datalad/datalad/issues/4777 +[iso1]: https://git-annex.branchable.com/projects/datalad/bugs-done/__34__357_out_of_984_tests_failed__34___on_NFS_lustre_mount/ +[glu1]: https://lists.gluster.org/pipermail/gluster-users/2015-May/021927.html +[gpf1]: https://github.com/datalad/datalad/issues/5589 +[gpf2]: https://handbook.datalad.org/en/inm7/intro/installation.html +[zfs1]: https://git-annex.branchable.com/todo/use_copy__95__file__95__range_for_get_and_copy/ + +### The pattern + +Four mechanisms account for nearly every report above: + +1. **SQLite cannot lock.** git-annex keeps its keys database in SQLite, + in WAL mode, which needs shared memory plus byte-range locks. Lustre, + CIFS, Gluster and WSL all fail this, and all produce the same + `SQLite3 returned ErrorIO` / `ErrorBusy` symptom. Joey's fix was + `annex.dbdir`, to move the database somewhere that works. +2. **POSIX record locks are missing or lying**, so git-annex falls back + to `annex.pidlock`. +3. **`link()` does not refuse an existing name** -- Lustre. That breaks + the pidlock fallback itself, which is why it is worth a check of its + own rather than being lumped in with (2). +4. **No symlinks / no exec bit / no ownership** -- the crippled-filesystem + path and adjusted branches. + +`bin/ci/fs-capabilities.sh` checks exactly these, plus the usual +suspects, and takes about a second. That matters for the recommendation +below: **a capability profile is a cheap approximation of a full `git +annex test` run**, and it can be collected for filesystems we cannot +afford to run the whole suite on. + +## Part 2 -- what can be bootstrapped, measured + +Every row below was produced by `bin/ci/probe-backend.sh ` on +a GitHub-hosted runner, not from documentation. `BOOTSTRAPPED` means it +mounted and was probed; `PARTIAL` means the interesting half is out of +reach on a runner; failures carry their reason. Times are the whole +bring-up: package install, mkfs or server start, mount. + +Run 5 of the probe workflow, ubuntu-24.04 unless noted, kernel +`6.17.0-1022-azure` (`6.8.0-1064-azure` on the 22.04 image). + +| Candidate | Verdict | Time | What the capability probe found | +| --- | --- | --- | --- | +| `gfs2` | **BOOTSTRAPPED** | 28s (61s on 22.04) | Everything green: symlinks, fcntl locks, SQLite WAL, xattrs, `link()` correctly refusing an existing name | +| `ocfs2` | **BOOTSTRAPPED** | 39s (52s on 22.04) | Same -- a full POSIX profile | +| `cifs` | **BOOTSTRAPPED** | 26s | **`sqlite-wal=no`, `sqlite-delete-mode=no`**, `symlink=no`, `fifo=no`, `exec-bit=no`, `perm-bits=no`, `case-sensitive=no` | +| `exfat` | **BOOTSTRAPPED** | 25s | Crippled as expected: no symlinks, hardlinks, fifos, exec bit, xattrs; case-insensitive; **rejects `:*?` in names** | +| `s3-rclone` | **BOOTSTRAPPED** | 46s | Crippled in the same shape as exfat -- but SQLite works | +| `sshfs` | **BOOTSTRAPPED** | 10s | **`hardlink-same-inode=no`** (link() succeeds, the link is invisible), no fifos, no unix sockets, no xattrs, and **1-second timestamp granularity** (1 distinct mtime across 5 rapid creates, vs 5 everywhere else) | +| `glusterfs` | **BOOTSTRAPPED** | 18s | Full POSIX profile through the FUSE client, SQLite WAL included -- see the note below, this does *not* refute the reported WAL failure | +| `zfs` | **BOOTSTRAPPED** | 13s | Full POSIX profile | +| `f2fs` | **BOOTSTRAPPED** | 39s | Full POSIX profile | +| `ntfs3` | **BOOTSTRAPPED** | 22s | Full POSIX profile -- the in-kernel driver carries symlinks and mode bits, unlike vfat | +| `nilfs2` | **BOOTSTRAPPED** | 27s | Full profile except `xattr-user=no` | +| `udf` | **BOOTSTRAPPED** | 9s | Full profile except xattrs and **255-character names** | +| `bcachefs` | **BOOTSTRAPPED** | 33s | Full POSIX profile | +| `overlay` | **BOOTSTRAPPED** | 0s | Full POSIX profile | +| `gocryptfs` | **BOOTSTRAPPED** | 9s | Full POSIX profile | +| `encfs` | **BOOTSTRAPPED** | 9s | Full profile except **255-character names** -- encryption inflates the name past the underlying limit | +| `vm-only` | *informational* | 0s | `/dev/kvm` **present**; `qemu-system-x86_64` not installed but one `apt install` away | +| `ecryptfs` | **NOT BOOTSTRAPPABLE** | 14s | `mount -t ecryptfs` refused with the key options the probe passes | +| `openafs` | **NOT BOOTSTRAPPABLE** | 295s | `openafs-modules-dkms` fails to build against `6.17.0-1022-azure` | +| `cephfs` | **NOT BOOTSTRAPPABLE** | 382s | The all-in-one demo container starts, the cluster never becomes responsive | +| `lustre-client` | **NOT BOOTSTRAPPABLE** | 15s / 0s | 22.04: no `lustre-client-modules-dkms` in the ubuntu2204 repo. 24.04: no ubuntu2404 client repo under `latest-release`, which is what the probe queries (see below) | + +Three of those measurements are worth pulling out, because they change +what the rows would be *for*: + +- **CIFS is both crippled and SQLite-hostile.** It is the only candidate + that fails `sqlite-wal` *and* `sqlite-delete-mode`, which is precisely + the [Samba report][smb1] and the same mechanism as the Lustre one -- + reproduced locally, in 26 seconds, with no HPC site involved. +- **sshfs quantises timestamps to a second.** Everything else on this + list resolves five rapid creates into five distinct mtimes; sshfs + resolves them into one. That is the shape of bug that makes git's + racy-timestamp handling matter, and no filesystem currently in the + matrix exhibits it. +- **sshfs creates hardlinks that cannot be seen as hardlinks.** `ln` + succeeds; both names then report distinct inodes and `nlink=1`, + because SFTP has no way to say otherwise. This one was found the + expensive way -- `hardlink=yes` called sshfs healthy, and it took a + real `git annex test` run to discover that `add` on an adjusted + unlocked branch fails on it. The probe now measures + `hardlink-same-inode` and `hardlink-nlink` directly, which turns a + 20-minute suite run into a 10-second answer. See GOTCHAS.md. +- **`ntfs3` is not vfat.** The in-kernel NTFS driver reports symlinks, + mode bits and xattrs, so it would *not* be a second crippled-filesystem + row -- it would be a row testing a driver users reach through WSL and + external drives, with POSIX mostly intact. Measured on a volume made by + `mkfs.ntfs` here, which is not the same thing as a Windows-formatted + disk mounted with `umask` defaults. + +**The glusterfs row does not contradict the Gluster report in Part 1.** +Part 1 cites a 2015 gluster-users thread where SQLite WAL fails; the +probe measures `sqlite-wal=yes`. Both are true of different things: the +probe stands up a *single-brick, replica-1, localhost* volume, which is +the cheapest thing that mounts, and WAL needs shared-memory mmap that a +one-brick local volume can satisfy. The report is from a multi-brick +deployment. Read the green row as "the bring-up works and a trivial +volume is POSIX-clean", not as "that report was wrong" -- reproducing it +needs a real multi-brick volume, which is out of scope here. + +Also worth recording as a negative: **every single-node filesystem here +passes `link-eexist`**, the POSIX rule Lustre is reported to break. The +check is cheap and stays in the probe, but nothing reachable from a +runner reproduces that particular Lustre behaviour -- if we want it +tested, it has to be Lustre. + +## Part 3 -- the two that were asked about + +### "GFS" is three different filesystems + +Worth disambiguating before deciding anything, because the three have +completely different answers: + +| Reading | What it is | Can we test it? | +| --- | --- | --- | +| **GFS2** | Red Hat's cluster filesystem, in the mainline kernel | **Yes, trivially.** `mkfs.gfs2 -p lock_nolock -j 1` needs no cluster stack at all, the module ships in `linux-modules-extra`, and the whole thing is a loop device. Measured at well under a minute end to end. | +| **GlusterFS** | userspace scale-out filesystem, FUSE client | **Yes.** One brick, one replica-1 volume, `mount -t glusterfs` back over the loopback -- all from Ubuntu packages. | +| **GPFS / IBM Storage Scale** | IBM's proprietary parallel filesystem | **Not in public CI.** See below. | + +Given the HPC context the question came from, GPFS is the likely +intended reading -- and it is the one we cannot do. But GFS2 is nearly +free and exercises *cluster-filesystem locking*, which is the property +that breaks git-annex on the parallel filesystems we cannot run. That +makes it a genuine, if partial, stand-in rather than a consolation prize. + +**GPFS specifically.** IBM ships a [Storage Scale Developer +Edition](https://www.ibm.com/docs/en/storage-scale/5.1.7?topic=overview-spectrum-scale-product-editions) +free for non-production use, capped at 12 TiB. It is a real option for a +*self-hosted* runner or a lab VM: it is a normal cluster install plus a +portability-layer kernel module built against the running kernel. It is +not an option for this repo's public CI -- the download is behind an IBM +ID, the license does not permit redistribution, and there is no apt repo +to point a workflow at. If GPFS coverage matters, the realistic path is +a self-hosted runner inside an institution that already licenses it. + +### Lustre + +**The client is the wall, and it is a hard one.** Lustre's client is an +out-of-tree kernel module. Whamcloud publishes Ubuntu packages, but as +modules *prebuilt per kernel* -- their ubuntu2204 index lists names like +[`lustre-client-modules-5.15.0-39-generic`][lpkg] -- and GitHub's runners +do not run those kernels -- measured, the runner is on +`6.8.0-1064-azure` (ubuntu-22.04 image) or `6.17.0-1022-azure` +(ubuntu-24.04). Also measured: `latest-release/ubuntu2204/client` has no +`lustre-client-modules-dkms` to build from source against it, and +`latest-release` -- the 2.15.x LTS line, which is what the probe queries +-- publishes no `ubuntu2404` client repo. + +Be careful how far that last point is taken: the 2.16.x *feature* line +does publish one (`lustre-2.16.0/ubuntu2404`, +`latest-feature-release/ubuntu2404/client`), and 2.16.0 lists Ubuntu +24.04 as a supported client platform. It does not rescue us, because +those modules are prebuilt for `6.8.0-35-generic` and the runner is on +an azure kernel -- so the conclusion is unchanged while the reason is +narrower than "no repo exists". Neither the prebuilt route nor the DKMS +route is open on a GitHub-hosted runner, and that is before the server +question. + +**The server is a second wall.** Lustre's ldiskfs OSD needs a *patched* +kernel; only the ZFS OSD runs on an unpatched one. So even with a +working client, a single-node Lustre means a ZFS-backed MGS/MDT/OST -- +which is a documented configuration ([ZFS OSD][zfsosd]), just not a small +one. + +**What does work, and is the recommended path:** Lustre's own test suite +ships `lustre/tests/llmount.sh` with a `cfg/local.sh` that builds a +complete single-node filesystem on **loopback devices**, no extra +hardware and no second machine ([Testing a Lustre filesystem][ltest]). +That is precisely an `eval-under` backend -- but it has to run somewhere +with a kernel Lustre supports. Three ways to get one, in increasing +order of effort: + +1. **A VM on the runner.** GitHub-hosted Linux runners *do* expose + `/dev/kvm` -- the `vm-only` probe reports it present, and the kernel + log shows `kvm_amd: Nested Virtualization enabled`. `qemu-system-x86_64` + is not preinstalled but is one `apt install` away. A Rocky/AlmaLinux 8 or 9 guest with + Whamcloud's repos, `llmount.sh` inside it, and the suite driven over + SSH is entirely feasible. Cost: several minutes of boot and install + per run, and a new `bin/eval-under-lustre` that is really "run this + inside a VM". The generic half of that (a `vm` backend that any + kernel-module filesystem could reuse -- Lustre today, GPFS on a + self-hosted runner tomorrow) is the part worth building. +2. **Extend the existing Vagrant VM.** `Vagrantfile` already provisions + a local box for BeeGFS work that CI cannot do. Adding a Lustre + flavour there gets a reproducible local target immediately, with no + CI cost, and is the cheapest way to find out what `git annex test` + does on Lustre. Prior art: the Lustre wiki's [Vagrant HPC storage + cluster][lvag] and [`dtrudg/vagrant-lustre-tutorial`][lvag2]. +3. **A real filesystem.** Amazon FSx for Lustre, or a self-hosted runner + at a site that already has Lustre. Highest fidelity, needs money or + institutional access, and cannot be a required check on a public PR. + +Recommendation: **(2) first** -- it costs a Vagrantfile stanza and +answers the actual question ("what breaks on Lustre?") for a human +sitting at a terminal. Promote to (1) only if the answer turns out to be +interesting enough to want a badge for. + +[lpkg]: https://downloads.whamcloud.com/public/lustre/lustre-2.15.1/ubuntu2204/client/Packages +[zfsosd]: https://wiki.lustre.org/ZFS_OSD +[ltest]: https://wiki.whamcloud.com/display/PUB/Testing+a+Lustre+filesystem +[lvag]: https://wiki.lustre.org/Create_a_Virtual_HPC_Storage_Cluster_with_Vagrant +[lvag2]: https://github.com/dtrudg/vagrant-lustre-tutorial + +## Part 4 -- recommendation + +**Tier 1 -- add as matrix rows.** Documented breakage, and they come up +on a stock runner in under a minute. + +- **`cifs`** (localhost Samba). The highest-value row on this list: it + reproduces the SQLite failure mode *directly* -- the probe measured + `sqlite-wal=no` and `sqlite-delete-mode=no` on it -- and SMB shares are + something ordinary users have, not just HPC sites. +- **`loop --fs gfs2`** and **`loop --fs ocfs2`**. Cluster-filesystem + locking, no cluster. Already implemented in `bin/eval-under-loop`. + +**Tier 2 -- each adds an axis the matrix does not have yet.** + +- **`sshfs`**, now the strongest candidate on this list: it is the only + one measured that breaks a documented git-annex operation outright + (any unlocked `git annex add`, via invisible hardlinks -- including but + not limited to an adjusted unlocked branch), *and* + it is the only one with a one-second clock. Backend implemented in a + separate change (`bin/eval-under-sshfs`); the measurements above come + from `bin/ci/probe-backend.sh`, which mounts sshfs itself. +- **`loop --fs exfat`**, the crippled row that is *not* vfat: it also + rejects `:`, `*` and `?` in filenames, which vfat-with-defaults does + not surface, and it is what is on every USB drive. +- **`ntfs3`**, which the probe shows is *not* a crippled filesystem -- + symlinks, mode bits and xattrs all work -- so it would test the + in-kernel NTFS driver users reach through WSL and external drives + rather than being a second vfat. It needs a small backend addition + first, though: there is no `mkfs.ntfs3` (the probe uses `mkfs.ntfs -F + -Q` and then mounts `-t ntfs3`), `install-backend.sh` has no ntfs entry + in its loop case, and `bin/eval-under-loop`'s `ntfs` branch mounts + without `-t ntfs3`, so with ntfs-3g installed it would land on the FUSE + driver instead of the kernel one. Note also that the profile was + measured on a Linux-created volume; a Windows-formatted disk mounted + with `umask` defaults looks considerably more crippled. +- **`s3-rclone`**, object storage seen as a filesystem: crippled in the + exfat shape but with SQLite working, which is a combination nothing + else on the list produces. + +**Tier 3 -- worth doing, not trivially.** `glusterfs` bootstraps in 18s +and shows a clean POSIX profile, so the interesting question there is +what `git annex test` does to it, not whether it runs -- but it needs +its own `bin/eval-under-glusterfs` rather than a loop flavour. `cephfs` +and Lustre both need more machine than a runner gives (see above). + +**Not feasible on GitHub-hosted runners.** GPFS (license), Lustre and +OpenAFS (out-of-tree modules that do not build against the runners' +Azure kernels). All three are feasible in a VM or on a self-hosted +runner, which is the same conclusion arrived at three different ways. + +### Promoting a candidate to a matrix row + +For anything the `loop` backend already handles, it is a data edit in +`evals/matrix.yaml` plus a regeneration: + +```yaml +# under `backends:`, next to the existing loop rows: + - backend: loop + version: gfs2 + label: Loop gfs2 + - backend: loop + version: ocfs2 + label: Loop ocfs2 +``` + +```bash +bin/ci/gen-readme-matrix.sh # refreshes the README badge grid +``` + +That is deliberately *not* done in this change: each row is four more +cells, four more badges and four more weekly runs against the shared +runner pool, and that is a call for whoever owns the CI budget. The +backend work is done; the matrix stays as it is until someone says go. + +A filesystem the loop backend cannot make (cifs, glusterfs, cephfs) +needs a `bin/eval-under-` first. `bin/ci/probe-backend.sh` already +contains a working bring-up sequence for each of them -- the `setup_*` +function is the backend, minus the flag handling. + +### The cheap target + +`bin/ci/fs-capabilities.sh` is now a target in its own right, +`capabilities`, alongside `git-annex`, `git`, `stress-ng` and +`pjdfstest`. It runs in about a second and answers "what does this +filesystem support?" for any backend, including the ones where the full +suite is too slow to run often. Where `pjdfstest` says which syscall +violates POSIX, this says which of the git-annex-specific mechanisms is +present. + +It is marked `on-demand: true` in `evals/matrix.yaml`, so it is +runnable but is not a matrix cell -- it adds no badges and no weekly +runs, and it is the default target of the reproduce workflow. Promoting +it to a scheduled column later is a one-line data edit: drop the flag. + +## Part 5 -- what this does not cover + +- **AFS.** No reports found either way, and the client half is now + measured as unavailable here: `openafs-modules-dkms` does not build + against the runner's `6.17.0-1022-azure` kernel. Same shape as Lustre, + same answer -- it needs a VM with a kernel the module supports. +- **9p and virtiofs** -- the WSL2 and VM-shared-folder paths, and a real + source of user reports. Both need a VM, so they land in the same + bucket as Lustre. +- **NFS variants.** Still the biggest gap, and it is about *options* + rather than filesystems: v3 vs v4.x, `actimeo=0`, `nolock`, `sync`. + See GOTCHAS.md, "Not yet covered". +- **eCryptfs.** Probed; `mount -t ecryptfs` refuses the key options the + probe passes, so the bring-up is a scripting problem rather than a + proven impossibility. Worth one more look, since it is the filesystem + behind Ubuntu's old encrypted-home feature and has a long history of + filename-length surprises -- and `encfs`, which *does* come up, already + shows that shape (`long-name-255=no`). +- **CephFS.** The demo container starts but the cluster never becomes + responsive inside the probe's budget on a 4-core runner. Not proven + impossible; proven not-cheap. diff --git a/GOTCHAS.md b/GOTCHAS.md index 06bba2b..e21bca3 100644 --- a/GOTCHAS.md +++ b/GOTCHAS.md @@ -23,11 +23,28 @@ A sparse backing image is `dd`'d, `mkfs.`'d, and loop-mounted. | Knob | Value | Why | | --- | --- | --- | -| `mkfs` options | none -- distro defaults | Whatever a user gets from `mkfs.ext4 /dev/sdX`, deliberately. | -| Image size | per target, `loop-size-mb` in `evals/matrix.yaml` | `git annex test` needs room for many small objects; the other three do not. | +| `mkfs` options | none -- distro defaults -- **except** the filesystems below | Whatever a user gets from `mkfs.ext4 /dev/sdX`, deliberately. | +| `mkfs.gfs2` | `-O -p lock_nolock -j 1 -J 8` | `lock_nolock` is GFS2's single-node lock module: no dlm, no corosync, no pacemaker. One journal because one node mounts it; `-J 8` because the 128MB default journal does not fit in a test-sized image. | +| `mkfs.ocfs2` | `-F -M local -N 1 -b 4K -C 8K --fs-features=local -q` | `-M local` is OCFS2's equivalent: a local mount needs no o2cb cluster stack. | +| Image size | per target, `loop-size-mb` in `evals/matrix.yaml`, raised to a per-filesystem floor | `git annex test` needs room for many small objects; the other three do not. A *default* `mkfs.gfs2` wants a 128MB journal, which will not fit in a 100MB image; with this backend's `-J 8` it does fit -- measured -- but 100MB then leaves `git annex test` almost no room, so gfs2 and ocfs2 are floored at 256MB (btrfs at 120MB). The floor is applied unconditionally and logged: `run-under.sh` always passes a `--size` chosen for the target, and no caller knows every filesystem's journal overhead. | | Mount (vfat, msdos, exfat, ntfs) | `-o uid=,gid=` | These filesystems store no ownership. Without `uid=`, everything belongs to root and an unprivileged wrapped command cannot write. | +| Mount (gfs2) | `-t gfs2 -o lockproto=lock_nolock`, then `chown` | Named again at mount time so an image labelled for a cluster still mounts single-node. | | Mount (everything else) | plain `mount`, then `chown ` on the mountpoint | ext4/xfs/btrfs carry real ownership; setting it once on the root is enough. | +**Single-node cluster filesystems are a deliberate approximation.** The +two vendors say different things about it, and neither says quite what an +earlier revision of this file claimed ("development-and-test-only"): Red +Hat does not support GFS2 as a single-node filesystem *at all*, outside +backup / secondary-site DR and existing single-node customers, and points +at a local filesystem instead; Oracle documents a local OCFS2 mount as a +supported configuration you can later migrate into a cluster, not as a +test-only mode. Either way, neither mode exercises the distributed lock +manager -- which is exactly the part a real GFS2 or OCFS2 deployment +would stress. What the row *does* buy is a cluster filesystem's on-disk +and VFS behaviour under git-annex, at loop-device cost. Read a green +cell as "nothing here is broken by the filesystem itself", not as "GFS2 +is fine in a cluster". + **The vfat consequence worth knowing.** `fmask`, `dmask` and `umask` are left at kernel defaults, so every file on the mount reads as mode `0755` and every directory as `0755`. `chmod` cannot change that -- vfat has one diff --git a/README.md b/README.md index 23f7932..cdcec95 100644 --- a/README.md +++ b/README.md @@ -14,6 +14,11 @@ itself is both backend- and suite-agnostic: new filesystems drop in as `bin/eval-under-` scripts, new suites as `bin/ci/target-.sh` (see below). +> **Which filesystem should be next?** [FILESYSTEMS.md](FILESYSTEMS.md) +> surveys the filesystems with documented git-annex / DataLad breakage +> against what can actually be stood up in CI -- the second half +> measured, not guessed, by the *Probe candidate filesystems* workflow. +> > **Read [GOTCHAS.md](GOTCHAS.md) before drawing conclusions from a red > cell.** It records the exact mkfs / mount / export settings each > backend uses -- results only mean something relative to those -- and @@ -196,6 +201,40 @@ four commits past it, and a `-dirty` suffix for uncommitted changes. An installed copy outside a checkout reports the `VERSION_FALLBACK` baked into `bin/eval-under`, bumped with each release tag. +## Reproducing a report + +The scheduled matrix answers "is this filesystem broken?" for the five +backends it covers. Support needs the other question answered: someone +reports a problem on a filesystem, possibly one with no cell, and you +want that combination running now with *their* mount options. + +Start with the capability profile -- seconds, and it usually explains +the failure before a suite is worth starting: + +```bash +sudo bin/eval-under nfs --set-home -- bin/ci/fs-capabilities.sh +``` + +Every line it prints maps to a class of reported bug (`sqlite-wal=no` -> +"SQLite3 returned ErrorIO"; `fcntl-lock=no` -> git-annex falls back to +`annex.pidlock`; `symlink=no` -> crippled filesystem, adjusted branch). +See FILESYSTEMS.md for which report each one came from. + +Then run the suite under the same backend: + +```bash +sudo bin/ci/run-under.sh nfs n/a git-annex +``` + +In CI, the **Reproduce (on demand)** workflow is the same thing from the +Actions tab: pick the backend, version, target, runner image, and any +backend flags (`--sync`, `--no-root-squash`, ...). It has +no badge and no schedule -- it exists to be run at someone, once. + +Targets available there include `capabilities`, which is not a matrix +column: it is the fast triage step above, wrapped so it can run under +any backend. + ## File layout | Path | Purpose | @@ -227,6 +266,12 @@ into `bin/eval-under`, bumped with each release tag. | `tests/test_*.py`, `tests/data/` | Unit tests of the results parsers and the known-issues classifier | | `.github/workflows/test.yaml` | The whole matrix: one `matrix` job, 20 `test` cells, one `publish` job | | `.github/workflows/checks.yaml` | `run-checks.sh` on every push and PR; minutes, no root, no mount | +| `.github/workflows/reproduce.yaml` | On-demand run of any backend x target, for support (no badge, no schedule) | +| `.github/workflows/probe-filesystems.yaml` | Runs the probes; a roll-up job renders the whole table | +| `bin/ci/probe-backend.sh` | Reconnaissance: try to stand up a candidate filesystem here, report what it got | +| `bin/ci/render-probe-table.sh` | Rolls those probe summaries up into the one table the workflow prints | +| `bin/ci/fs-capabilities.sh` | What a mounted filesystem supports, in the dimensions git-annex trips over | +| `FILESYSTEMS.md` | Survey: reported git-annex / DataLad breakage vs. what can be bootstrapped here | | `drafts/git-annex-test-beegfs.yaml` | Copy-target workflow for `con/git-annex` (external PR target) | ## Local iteration (VM) diff --git a/bin/ci/evals.py b/bin/ci/evals.py index 38121fd..b0920d0 100644 --- a/bin/ci/evals.py +++ b/bin/ci/evals.py @@ -32,11 +32,17 @@ def load_matrix(path: Path = MATRIX_FILE) -> dict: def matrix_cells(m: dict) -> dict[str, dict]: - """slug -> cell metadata, in matrix (row, column) order.""" + """slug -> cell metadata, in matrix (row, column) order. + + An `on-demand` target is runnable but is not a cell: it has no + scheduled run, so it has no status, and including it here would + publish a permanently-"unknown" badge and a column that can never + fill in. + """ cells = {} for b in m["backends"]: bslug = backend_slug(b["backend"], b["version"]) - for t in m["targets"]: + for t in (t for t in m["targets"] if not t.get("on-demand")): cells[f"{bslug}-{t['name']}"] = { "backend": b["backend"], "version": b["version"], diff --git a/bin/ci/fs-capabilities.sh b/bin/ci/fs-capabilities.sh new file mode 100755 index 0000000..d17b577 --- /dev/null +++ b/bin/ci/fs-capabilities.sh @@ -0,0 +1,234 @@ +#!/bin/bash +# SPDX-FileCopyrightText: 2026 Yaroslav Halchenko +# SPDX-License-Identifier: MIT +# +# Generated with Claude Code +# +# Report the POSIX-ish capabilities of a directory's filesystem, in the +# specific dimensions that git-annex / DataLad are known to trip over. +# +# Every check here corresponds to a real reported failure -- see +# FILESYSTEMS.md for the issue each one maps to. The point is that +# a capability profile is orders of magnitude cheaper than `git annex +# test` and usually explains its result: +# +# sqlite-wal=no -> "SQLite3 returned ErrorIO" (Lustre, WSL, CIFS) +# fcntl-lock=no -> git-annex falls back to annex.pidlock (NFS) +# link-eexist=no -> pidlock is unsafe: two files, one name (Lustre) +# symlink=no -> crippled filesystem, adjusted branch (vfat) +# hardlink-same-inode=no -> "failed to link to annex" (sshfs): link() +# works but the link is not observable as one +# exec-bit=no -> vfat/ntfs: every file looks executable +# +# usage: +# bin/ci/fs-capabilities.sh [DIR] +# +# DIR directory to probe (default: $EVAL_UNDER_MOUNT, else $TMPDIR, +# else the current directory) +# +# Output is one `key=value` per line on stdout, plus a human-readable +# `fstype`/`mount` header on stderr. Always exits 0 unless DIR is +# unusable: an unsupported feature is a finding, not an error. + +set -uo pipefail + +DIR="${1:-${EVAL_UNDER_MOUNT:-${TMPDIR:-.}}}" + +[ -d "$DIR" ] || { echo "not a directory: $DIR" >&2; exit 2; } + +work="$(mktemp -d "$DIR/fscaps-XXXXXX")" || { + echo "cannot create a working directory under $DIR" >&2; exit 2; } +trap 'rm -rf "$work" 2>/dev/null || true' EXIT + +say() { printf '%s=%s\n' "$1" "$2"; } + +# Several checks below are python3 one-liners. Without python3 they would +# each exit non-zero and be reported as a missing *filesystem feature* -- +# including sqlite-wal, the most predictive value in this whole profile -- +# so a profile pasted into a bug report would blame the filesystem for a +# missing interpreter. Refuse to produce a half-profile instead. +command -v python3 >/dev/null 2>&1 || { + echo "python3 is required: without it, checks that need it would be" >&2 + echo "reported as unsupported filesystem features rather than as a" >&2 + echo "missing tool." >&2 + exit 2 +} + +# Run a check in a subshell; "yes" if it exits 0, "no" otherwise. Output +# is swallowed -- the verdict is the value. +# +# Exit 127 is kept distinct: a check whose helper is absent reports +# "unknown", never "no", for the same reason as the python3 preflight +# above. +check() { + local key="$1"; shift + if ( "$@" ) >/dev/null 2>&1; then + say "$key" yes + elif [ "$?" = 127 ]; then + say "$key" unknown + else + say "$key" no + fi +} + +{ + echo "dir: $DIR" + echo "fstype: $(stat -f -c %T "$DIR" 2>/dev/null || echo unknown)" + findmnt -n -o SOURCE,FSTYPE,OPTIONS --target "$DIR" 2>/dev/null || true +} >&2 + +say fstype "$(stat -f -c %T "$DIR" 2>/dev/null || echo unknown)" + +c_symlink() { ln -s target "$work/sl" && [ -L "$work/sl" ]; } +c_hardlink() { : > "$work/hl-a" && ln "$work/hl-a" "$work/hl-b"; } + +# `ln` succeeding is not the same as the link being observable. sshfs +# creates the link over SFTP but reports a distinct inode and nlink=1 for +# each name -- so the two names are indistinguishable from two copies. +# git-annex's add (adjusted unlocked branch) hardlinks the file into +# .git/annex/objects and then verifies it, and git's own local clone does +# the same check, so both fail on such a filesystem while `ln` looks fine. +# These two checks are what predict that, and c_hardlink alone does not. +c_hardlink_same_inode() { + : > "$work/hli-a" && ln "$work/hli-a" "$work/hli-b" || return 1 + local ia ib + ia="$(stat -c %i "$work/hli-a")" && ib="$(stat -c %i "$work/hli-b")" || return 1 + [ "$ia" = "$ib" ] +} +c_hardlink_nlink() { + : > "$work/hln-a" && ln "$work/hln-a" "$work/hln-b" || return 1 + [ "$(stat -c %h "$work/hln-a")" = 2 ] +} +c_fifo() { mkfifo "$work/fifo"; } +c_unix_socket() { python3 -c ' +import socket, sys, os +s = socket.socket(socket.AF_UNIX) +s.bind(os.path.join(sys.argv[1], "sock")) +' "$work"; } + +# vfat & friends: mode bits are synthesised from mount options, so a +# freshly-created 0644 file reads back as executable and chmod is a no-op. +c_exec_bit() { + : > "$work/x"; chmod 644 "$work/x"; [ ! -x "$work/x" ] || return 1 + chmod 755 "$work/x"; [ -x "$work/x" ] +} +c_perm_bits() { + : > "$work/p"; chmod 600 "$work/p" + [ "$(stat -c %a "$work/p")" = "600" ] +} + +# POSIX: link() must fail with EEXIST if newpath exists. Lustre has been +# observed to succeed here and end up with two directory entries of the +# same name -- which is what makes annex.pidlock unsafe on it. +c_link_eexist() { + : > "$work/le-a"; : > "$work/le-b" + ! ln "$work/le-a" "$work/le-b" 2>/dev/null +} +c_rename_over() { + echo a > "$work/ro-a"; echo b > "$work/ro-b" + mv -f "$work/ro-a" "$work/ro-b" && [ "$(cat "$work/ro-b")" = a ] +} +# Deleting an open file: NFS silly-renames it (.nfsXXXX), most local +# filesystems just unlink it. +c_unlink_open() { python3 -c ' +import os, sys +p = os.path.join(sys.argv[1], "uo") +fh = open(p, "w") +os.unlink(p) +fh.write("still writable") +fh.close() +sys.exit(0 if not os.path.exists(p) else 1) +' "$work"; } + +c_fcntl_lock() { python3 -c ' +import fcntl, os, sys +fh = open(os.path.join(sys.argv[1], "fl"), "w") +fcntl.lockf(fh, fcntl.LOCK_EX | fcntl.LOCK_NB) +fcntl.lockf(fh, fcntl.LOCK_UN) +' "$work"; } +c_flock() { python3 -c ' +import fcntl, os, sys +fh = open(os.path.join(sys.argv[1], "fk"), "w") +fcntl.flock(fh, fcntl.LOCK_EX | fcntl.LOCK_NB) +fcntl.flock(fh, fcntl.LOCK_UN) +' "$work"; } + +# The single most predictive check for git-annex: its keys database is +# SQLite in WAL mode, and WAL needs shared memory + byte-range locks. +c_sqlite_wal() { python3 -c ' +import os, sqlite3, sys +db = os.path.join(sys.argv[1], "wal.db") +con = sqlite3.connect(db) +mode = con.execute("PRAGMA journal_mode=WAL;").fetchone()[0] +if mode.lower() != "wal": + sys.exit(1) +con.execute("CREATE TABLE t (a int);") +con.execute("INSERT INTO t VALUES (1);") +con.commit() +con.close() +sqlite3.connect(db).execute("SELECT * FROM t;").fetchall() +' "$work"; } +c_sqlite_delete_mode() { python3 -c ' +import os, sqlite3, sys +con = sqlite3.connect(os.path.join(sys.argv[1], "del.db")) +con.execute("PRAGMA journal_mode=DELETE;") +con.execute("CREATE TABLE t (a int);") +con.execute("INSERT INTO t VALUES (1);") +con.commit() +' "$work"; } + +c_xattr_user() { python3 -c ' +import os, sys +p = os.path.join(sys.argv[1], "xa") +open(p, "w").close() +os.setxattr(p, "user.eval_under", b"1") +os.getxattr(p, "user.eval_under") +' "$work"; } +c_chown() { : > "$work/co" && chown 1:1 "$work/co"; } +c_case_sensitive() { + : > "$work/CaseA"; [ ! -e "$work/casea" ] +} +c_long_name() { : > "$work/$(printf 'n%.0s' $(seq 1 255))"; } +# git-annex writes keys containing characters some filesystems reject. +c_special_chars() { : > "$work/a:b*c?d"; } +c_trailing_space() { : > "$work/trailing "; } + +check symlink c_symlink +check hardlink c_hardlink +check hardlink-same-inode c_hardlink_same_inode +check hardlink-nlink c_hardlink_nlink +check fifo c_fifo +check unix-socket c_unix_socket +check exec-bit c_exec_bit +check perm-bits c_perm_bits +check link-eexist c_link_eexist +check rename-over c_rename_over +check unlink-open c_unlink_open +check fcntl-lock c_fcntl_lock +check flock c_flock +check sqlite-wal c_sqlite_wal +check sqlite-delete-mode c_sqlite_delete_mode +check xattr-user c_xattr_user +if [ "$(id -u)" = 0 ]; then + check chown c_chown +else + # chown(2) is restricted to root regardless of the filesystem, so a + # "no" here would say nothing about the filesystem. + say chown "n/a-not-root" +fi +check case-sensitive c_case_sensitive +check long-name-255 c_long_name +check special-chars c_special_chars +check trailing-space c_trailing_space + +# Timestamp granularity: how many distinct mtimes we see across rapid +# writes. vfat rounds to 2s, which is enough to hide a modification from +# a stat-based dirty check. +mtimes="$(for i in 1 2 3 4 5; do + : > "$work/ts$i" + stat -c %.9Y "$work/ts$i" 2>/dev/null || stat -c %Y "$work/ts$i" + sleep 0.05 +done | sort -u | wc -l)" +say mtime-distinct-of-5 "$mtimes" + +df -P -k "$DIR" | awk 'NR==2 {print "free-kb=" $4}' diff --git a/bin/ci/install-backend.sh b/bin/ci/install-backend.sh index 6c5ee54..de6c2d0 100755 --- a/bin/ci/install-backend.sh +++ b/bin/ci/install-backend.sh @@ -79,11 +79,30 @@ install_loop() { xfs) pkg=xfsprogs ;; btrfs) pkg=btrfs-progs ;; ext2|ext3|ext4) pkg=e2fsprogs ;; # usually preinstalled + gfs2) pkg=gfs2-utils ;; + ocfs2) pkg=ocfs2-tools ;; + f2fs) pkg=f2fs-tools ;; + exfat) pkg=exfatprogs ;; + nilfs2) pkg=nilfs-tools ;; + bcachefs) pkg=bcachefs-tools ;; *) echo "unknown loop filesystem: $VERSION" >&2; exit 1 ;; esac apt_update apt_install "$pkg" command -v "mkfs.$VERSION" + + # gfs2 and ocfs2 ship in linux-modules-extra rather than the base + # kernel package. It is present on GitHub-hosted runners, but say so + # out loud here rather than discovering it as an "unknown filesystem + # type" at mount time. + case "$VERSION" in + gfs2|ocfs2) + if ! sudo modprobe "$VERSION"; then + apt_install "linux-modules-extra-$(uname -r)" + sudo modprobe "$VERSION" + fi + ;; + esac } case "$BACKEND" in diff --git a/bin/ci/install-target.sh b/bin/ci/install-target.sh index 508eebc..4e9c16b 100755 --- a/bin/ci/install-target.sh +++ b/bin/ci/install-target.sh @@ -36,7 +36,8 @@ here="$(cd "$(dirname "$0")" && pwd)" TARGET="${1:?target required (git-annex|git|stress-ng|pjdfstest)}" target_known "$TARGET" || { - echo "unknown target: $TARGET (expected: ${EVAL_UNDER_TARGETS[*]})" >&2 + echo "unknown target: $TARGET (expected: ${EVAL_UNDER_TARGETS[*]}" \ + "${EVAL_UNDER_ONDEMAND_TARGETS[*]})" >&2 exit 1 } @@ -141,4 +142,7 @@ case "$TARGET" in git) install_git ;; stress-ng) install_stress_ng ;; pjdfstest) install_pjdfstest ;; + # Needs nothing installed: it is bin/ci/fs-capabilities.sh, which is + # in this repo and uses only coreutils + sqlite3 if present. + capabilities) echo "I: $TARGET needs no runner-side install" ;; esac diff --git a/bin/ci/matrix.sh b/bin/ci/matrix.sh index 3700166..bcfe1e0 100755 --- a/bin/ci/matrix.sh +++ b/bin/ci/matrix.sh @@ -46,11 +46,22 @@ backends = d["backends"] if not targets or not backends: sys.exit("matrix.yaml: empty targets or backends") +# A target flagged `on-demand: true` is runnable but is not a matrix +# cell: it exists for reproduction (.github/workflows/reproduce.yaml), +# not for a scheduled badge. Splitting the list here is what keeps +# matrix-json.sh -- which iterates EVAL_UNDER_TARGETS -- from generating +# cells for it, without every consumer needing to know about the flag. +scheduled = [t for t in targets if not t.get("on-demand")] +ondemand = [t for t in targets if t.get("on-demand")] +if not scheduled: + sys.exit("matrix.yaml: every target is on-demand; no cells to run") + def q(v): return shlex.quote(str(v)) out = [] -out.append("EVAL_UNDER_TARGETS=(%s)" % " ".join(q(t["name"]) for t in targets)) +out.append("EVAL_UNDER_TARGETS=(%s)" % " ".join(q(t["name"]) for t in scheduled)) +out.append("EVAL_UNDER_ONDEMAND_TARGETS=(%s)" % " ".join(q(t["name"]) for t in ondemand)) out.append("EVAL_UNDER_BACKENDS=(%s)" % " ".join( q("%s|%s|%s" % (b["backend"], b["version"], b["label"])) for b in backends)) @@ -112,9 +123,12 @@ cell_slug() { echo "$(backend_slug "$backend" "$version")-$target" } +# True for anything runnable, scheduled or on-demand. Callers that must +# enumerate *cells* (matrix-json.sh, gen-readme-matrix.sh) iterate +# EVAL_UNDER_TARGETS instead, which holds the scheduled ones only. target_known() { local t - for t in "${EVAL_UNDER_TARGETS[@]}"; do + for t in "${EVAL_UNDER_TARGETS[@]}" "${EVAL_UNDER_ONDEMAND_TARGETS[@]}"; do [ "$t" = "$1" ] && return 0 done return 1 diff --git a/bin/ci/probe-backend.sh b/bin/ci/probe-backend.sh new file mode 100755 index 0000000..4f570dc --- /dev/null +++ b/bin/ci/probe-backend.sh @@ -0,0 +1,529 @@ +#!/bin/bash +# SPDX-FileCopyrightText: 2026 Yaroslav Halchenko +# SPDX-License-Identifier: MIT +# +# Generated with Claude Code +# +# Answer one question empirically, for one candidate filesystem: +# +# can a throwaway instance of it be stood up on a stock GitHub-hosted +# Ubuntu runner, and if so, what does it actually support? +# +# This is *reconnaissance*, not a backend. A candidate that comes up +# green here is a candidate for a real `bin/eval-under-`; one that +# comes up red tells us why (no kernel module / no package / needs a +# cluster / needs a VM) without anyone having to write the backend +# first. See FILESYSTEMS.md for the survey this feeds. +# +# usage: +# bin/ci/probe-backend.sh CANDIDATE +# bin/ci/probe-backend.sh --list +# +# CANDIDATE one of the names printed by --list +# +# env: +# EVAL_UNDER_PROBE_DIR scratch directory (default /var/tmp/eu-probe) +# EVAL_UNDER_PROBE_KEEP non-empty: skip teardown (local debugging) +# +# Exit status is the finding: 0 = the filesystem mounted and was +# probed, 1 = it could not be brought up here. Either way the log says +# how far it got, and a one-line verdict lands in $GITHUB_STEP_SUMMARY. + +# Deliberately no `-e`: a probe that dies on the first failed command +# reports nothing. Failures are caught per step and turned into a +# verdict instead. +set -uo pipefail +export DEBIAN_FRONTEND=noninteractive + +HERE="$(cd "$(dirname "$0")" && pwd)" + +CANDIDATES=( + gfs2 ocfs2 f2fs exfat ntfs3 udf nilfs2 bcachefs zfs overlay + glusterfs cephfs cifs sshfs gocryptfs encfs ecryptfs s3-rclone + lustre-client openafs vm-only +) + +usage() { + cat <&2 + exit 2 +} + +APT_LOCK_TIMEOUT=(-o "DPkg::Lock::Timeout=120") +apt_update() { sudo apt-get "${APT_LOCK_TIMEOUT[@]}" update -qq; } +apt_install() { + apt_update || true + sudo apt-get "${APT_LOCK_TIMEOUT[@]}" install -y --no-install-recommends "$@" +} + +# Deferred teardown: commands are run in reverse registration order. +# +# Entries are eval'd, so every interpolated path has to arrive already +# quoted -- an EVAL_UNDER_PROBE_DIR containing a space used to split into +# two arguments and leave the mount behind while the probe still reported +# success. defer() quotes its arguments itself: pass the command as +# separate words, not as one pre-built string. +CLEANUPS=() +defer() { + local quoted + quoted="$(printf '%q ' "$@")" + CLEANUPS+=("$quoted") +} +# For the handful of cleanups that genuinely need shell syntax (||, ;). +# The caller is then responsible for quoting inside the string. +defer_shell() { CLEANUPS+=("$1"); } +# shellcheck disable=SC2317 # invoked from the EXIT trap, not inline +run_cleanups() { + [ -n "${EVAL_UNDER_PROBE_KEEP:-}" ] && { echo "I: --keep: leaving $MNT"; return; } + local i + for (( i=${#CLEANUPS[@]}-1; i>=0; i-- )); do + echo "I: cleanup: ${CLEANUPS[$i]}" + # shellcheck disable=SC2086 # cleanup entries are our own literals + eval "${CLEANUPS[$i]}" || true + done +} +trap run_cleanups EXIT + +# Fail a probe with a reason that ends up in the summary line. +give_up() { NOTE="$*"; echo "E: $*" >&2; return 1; } + +# Load a kernel module, installing the runner's modules-extra package +# first if the module is not there at all. On GitHub-hosted runners the +# kernel is linux-azure and several filesystems live in +# linux-modules-extra-, which is *not* always preinstalled. +need_module() { + local mod="$1" + if sudo modprobe "$mod" 2>/dev/null; then return 0; fi + echo "I: modprobe $mod failed, trying linux-modules-extra-$(uname -r)" + apt_install "linux-modules-extra-$(uname -r)" || true + sudo modprobe "$mod" 2>/dev/null || give_up "no kernel module: $mod" +} + +# A loop device backed by a sparse file. Sets $LOOP rather than echoing +# it: command substitution would run this in a subshell and the `defer` +# registration would be lost with it. +LOOP="" +make_loop() { + local size="${1:-1G}" img="$PROBE_DIR/img.$$" + truncate -s "$size" "$img" || return 1 + LOOP="$(sudo losetup --find --show "$img")" || return 1 + defer_shell "sudo losetup -d '$LOOP'; rm -f '$img'" +} + +mount_and_own() { + sudo mount "$@" || return 1 + defer_shell "sudo umount -l '$MNT'" + sudo chown "$(id -u):$(id -g)" "$MNT" 2>/dev/null || true +} + +# ---------------------------------------------------------------- probes + +# GFS2 and OCFS2 are cluster filesystems, but both have a single-node +# locking mode meant for exactly this: no corosync/dlm, no pacemaker, +# just a block device. Red Hat and Oracle both document it as +# test/development-only, which is what we are. +setup_gfs2() { + apt_install gfs2-utils || give_up "gfs2-utils not installable" || return 1 + need_module gfs2 || return 1 + local loop; make_loop 1G || return 1; loop="$LOOP" + sudo mkfs.gfs2 -O -p lock_nolock -j 1 "$loop" || give_up "mkfs.gfs2 failed" || return 1 + mount_and_own -t gfs2 -o lockproto=lock_nolock "$loop" "$MNT" +} + +setup_ocfs2() { + apt_install ocfs2-tools || give_up "ocfs2-tools not installable" || return 1 + need_module ocfs2 || return 1 + local loop; make_loop 1G || return 1; loop="$LOOP" + sudo mkfs.ocfs2 -F -b 4K -C 32K -N 1 -M local "$loop" \ + || give_up "mkfs.ocfs2 failed" || return 1 + mount_and_own -t ocfs2 "$loop" "$MNT" +} + +# Plain loop filesystems the existing `loop` backend does not cover yet. +setup_loopfs() { + local fs="$1" pkg="$2" mkfs_opts="$3" mount_opts="${4:-}" + apt_install "$pkg" || give_up "$pkg not installable" || return 1 + need_module "$fs" || return 1 + local loop; make_loop 1G || return 1; loop="$LOOP" + # shellcheck disable=SC2086 # mkfs_opts is a deliberate word list + sudo "mkfs.$fs" $mkfs_opts "$loop" || give_up "mkfs.$fs failed" || return 1 + if [ -n "$mount_opts" ]; then + mount_and_own -t "$fs" -o "$mount_opts" "$loop" "$MNT" + else + mount_and_own -t "$fs" "$loop" "$MNT" + fi +} + +setup_f2fs() { setup_loopfs f2fs f2fs-tools "-f"; } +setup_exfat() { setup_loopfs exfat exfatprogs "" "uid=$(id -u),gid=$(id -g)"; } +setup_nilfs2() { setup_loopfs nilfs2 nilfs-tools "-f"; } +setup_bcachefs() { setup_loopfs bcachefs bcachefs-tools "-f"; } + +setup_ntfs3() { + apt_install ntfs-3g || give_up "ntfs-3g not installable" || return 1 + need_module ntfs3 || return 1 + local loop; make_loop 1G || return 1; loop="$LOOP" + sudo mkfs.ntfs -F -Q "$loop" || give_up "mkfs.ntfs failed" || return 1 + mount_and_own -t ntfs3 -o "uid=$(id -u),gid=$(id -g)" "$loop" "$MNT" +} + +setup_udf() { + apt_install udftools || give_up "udftools not installable" || return 1 + need_module udf || return 1 + local loop; make_loop 1G || return 1; loop="$LOOP" + sudo mkudffs --media-type=hd "$loop" || give_up "mkudffs failed" || return 1 + mount_and_own -t udf "$loop" "$MNT" +} + +setup_zfs() { + apt_install zfsutils-linux || give_up "zfsutils-linux not installable" || return 1 + need_module zfs || return 1 + local img="$PROBE_DIR/zfs.img" + truncate -s 2G "$img" || return 1 + sudo zpool create -f -m "$MNT" eu_probe "$img" || give_up "zpool create failed" || return 1 + defer_shell "sudo zpool destroy eu_probe; rm -f '$img'" + sudo chown "$(id -u):$(id -g)" "$MNT" +} + +# The filesystem every containerised CI job is already sitting on. +setup_overlay() { + local base="$PROBE_DIR/ovl" + mkdir -p "$base"/{lower,upper,work} + mount_and_own -t overlay overlay \ + -o "lowerdir=$base/lower,upperdir=$base/upper,workdir=$base/work" "$MNT" +} + +# Single-node Gluster: one brick, one replica-1 volume, FUSE-mounted +# back over localhost. `force` is needed because the brick is on the +# root filesystem, which Gluster otherwise refuses. +setup_glusterfs() { + apt_install glusterfs-server || give_up "glusterfs-server not installable" || return 1 + sudo systemctl start glusterd || sudo glusterd || give_up "glusterd will not start" || return 1 + defer_shell "sudo systemctl stop glusterd || true" + sudo systemctl --no-pager status glusterd 2>&1 | head -5 || true + local brick="$PROBE_DIR/brick" host + # Gluster records the brick's host in the volfile and insists it be a + # peer; the node's own hostname is always one, 127.0.0.1 need not be. + host="$(hostname -s)" + sudo mkdir -p "$brick" + sudo gluster --mode=script volume create eu_probe "$host:$brick" force \ + || give_up "volume create failed (see output above)" || return 1 + defer_shell "sudo gluster --mode=script volume stop eu_probe; sudo gluster --mode=script volume delete eu_probe" + sudo gluster --mode=script volume start eu_probe || give_up "volume start failed" || return 1 + sudo gluster volume info eu_probe || true + mount_and_own -t glusterfs "$host:/eu_probe" "$MNT" +} + +# CephFS via the upstream all-in-one demo container: a MON/MGR/OSD/MDS +# in one image, then the in-kernel client. Heavier than the others but +# still a single `docker run`. +setup_cephfs() { + command -v docker >/dev/null || give_up "no docker" || return 1 + apt_install ceph-common ceph-fuse || give_up "ceph-common not installable" || return 1 + # `docker run -d` pulls first, and the demo image is about a gigabyte. + timeout 420 sudo docker run -d --name eu-ceph --net=host \ + -e MON_IP=127.0.0.1 -e CEPH_PUBLIC_NETWORK=127.0.0.1/32 \ + -e CEPH_DEMO_UID=eu -e DEMO_DAEMONS="mon,mgr,osd,mds" \ + -v /etc/ceph:/etc/ceph -v /var/lib/ceph:/var/lib/ceph \ + quay.io/ceph/demo || give_up "ceph demo container did not start in 7min" || return 1 + defer_shell "sudo docker rm -f eu-ceph; sudo rm -rf /etc/ceph/* /var/lib/ceph/*" + # Every ceph client call needs its own timeout. With no reachable mon, + # `ceph -s` blocks for minutes on its internal retry loop rather than + # failing -- which is what turned this probe into a 25-minute job that + # held the roll-up hostage. + local i + for i in $(seq 1 24); do + [ -f /etc/ceph/ceph.conf ] \ + && timeout 15 sudo ceph --connect-timeout 10 -s >/dev/null 2>&1 && break + sleep 5 + # ~2 minutes of polling on top of the pull. Longer than that and + # the answer is "not on a runner", which is a verdict. + [ "$i" = 24 ] && { sudo docker logs --tail 30 eu-ceph 2>&1 || true + give_up "ceph cluster never became responsive"; return 1; } + done + timeout 20 sudo ceph -s || true + timeout 60 sudo ceph-fuse "$MNT" || give_up "ceph-fuse mount failed" || return 1 + defer_shell "sudo umount -l '$MNT'" + sudo chown "$(id -u):$(id -g)" "$MNT" || true +} + +# SMB against a local Samba, i.e. what a user with a NAS or a Windows +# share actually has. Mounted with kernel defaults: no `nobrl`, no +# `noserverino` -- those are exactly the workarounds we would want a red +# cell to motivate. +setup_cifs() { + apt_install samba cifs-utils || give_up "samba not installable" || return 1 + local share="$PROBE_DIR/share" + sudo mkdir -p "$share" + sudo chmod 777 "$share" + # Write the share to its own include file and register the include, + # so teardown can remove the share without rewriting smb.conf. The + # previous version appended the share to smb.conf itself and only + # stopped smbd, so the next `systemctl start smbd` -- or a reboot -- + # re-exposed a mode-777 `force user = root` share, on every + # interface, together with a known SMB password for root. + local inc=/etc/samba/eu-probe.conf + sudo tee "$inc" >/dev/null </dev/null + defer sudo rm -f "$inc" + defer_shell "sudo sed -i '\\# include = $inc#d' /etc/samba/smb.conf || true" + + printf 'eupass\neupass\n' | sudo smbpasswd -s -a root || give_up "smbpasswd failed" || return 1 + # Take the password back out of the passdb, not just off the wire. + defer_shell "sudo smbpasswd -x root >/dev/null 2>&1 || true" + sudo systemctl restart smbd || give_up "smbd will not start" || return 1 + defer_shell "sudo systemctl stop smbd || true" + mount_and_own -t cifs "//127.0.0.1/euprobe" \ + -o "username=root,password=eupass,uid=$(id -u),gid=$(id -g),vers=3.1.1" "$MNT" +} + +setup_sshfs() { + apt_install sshfs openssh-server || give_up "sshfs not installable" || return 1 + sudo systemctl start ssh || sudo systemctl start sshd || give_up "no sshd" || return 1 + local key="$HOME/.ssh/eu_probe" + mkdir -p "$HOME/.ssh" + [ -f "$key" ] || ssh-keygen -t ed25519 -N '' -f "$key" -q -C eval-under-probe + # Appending to the real ~/.ssh/authorized_keys and leaving it there + # is what bin/eval-under-sshfs exists to avoid; the probe has no + # business being less careful. Register the removal of exactly the + # line we add before adding it. + cat "$key.pub" >> "$HOME/.ssh/authorized_keys" + chmod 600 "$HOME/.ssh/authorized_keys" + defer_shell "sed -i '/eval-under-probe/d' '$HOME/.ssh/authorized_keys' || true" + defer rm -f "$key" "$key.pub" + local backing="$PROBE_DIR/sshfs-backing" + mkdir -p "$backing" + sshfs -o "IdentityFile=$key,StrictHostKeyChecking=no,UserKnownHostsFile=/dev/null" \ + "$(id -un)@127.0.0.1:$backing" "$MNT" || give_up "sshfs mount failed" || return 1 + defer_shell "fusermount3 -u '$MNT' || fusermount -u '$MNT' || sudo umount -l '$MNT'" +} + +setup_gocryptfs() { + apt_install gocryptfs || give_up "gocryptfs not installable" || return 1 + local cipher="$PROBE_DIR/cipher" pw="$PROBE_DIR/pw" + mkdir -p "$cipher" + echo eupassphrase > "$pw" + gocryptfs -init -passfile "$pw" -q "$cipher" || give_up "gocryptfs init failed" || return 1 + gocryptfs -passfile "$pw" -q "$cipher" "$MNT" || give_up "gocryptfs mount failed" || return 1 + defer_shell "fusermount3 -u '$MNT' || fusermount -u '$MNT' || sudo umount -l '$MNT'" +} + +setup_encfs() { + apt_install encfs || give_up "encfs not installable" || return 1 + local cipher="$PROBE_DIR/encfs-cipher" + mkdir -p "$cipher" + echo eupassphrase | encfs --standard --stdinpass "$cipher" "$MNT" \ + || give_up "encfs mount failed" || return 1 + defer_shell "fusermount3 -u '$MNT' || fusermount -u '$MNT' || sudo umount -l '$MNT'" +} + +setup_ecryptfs() { + apt_install ecryptfs-utils keyutils || give_up "ecryptfs-utils not installable" || return 1 + need_module ecryptfs || return 1 + local lower="$PROBE_DIR/ecryptfs-lower" sig + sudo mkdir -p "$lower" + local sigout + sigout="$(printf 'eupassphrase' | sudo ecryptfs-add-passphrase --fnek 2>&1)" + echo "I: ecryptfs-add-passphrase said: $sigout" + sig="$(echo "$sigout" | sed -n 's/.*\[\([0-9a-f]\{16\}\)\].*/\1/p' | head -1)" + [ -n "$sig" ] || { give_up "ecryptfs-add-passphrase produced no signature"; return 1; } + mount_and_own -t ecryptfs "$lower" "$MNT" \ + -o "key=passphrase:passphrase_passwd=eupassphrase,ecryptfs_cipher=aes,ecryptfs_key_bytes=16,ecryptfs_sig=$sig,ecryptfs_unlink_sigs,no_sig_cache=yes" +} + +# S3 seen as a filesystem: MinIO in a container, rclone's FUSE mount on +# top. This is the "my data lives in object storage" shape, which git +# and git-annex hate for reasons worth measuring rather than assuming. +setup_s3_rclone() { + command -v docker >/dev/null || give_up "no docker" || return 1 + apt_install rclone || give_up "rclone not installable" || return 1 + sudo docker run -d --name eu-minio -p 9000:9000 \ + -e MINIO_ROOT_USER=euprobe -e MINIO_ROOT_PASSWORD=euprobe123 \ + quay.io/minio/minio server /data || give_up "minio did not start" || return 1 + defer_shell "sudo docker rm -f eu-minio" + local i + for i in $(seq 1 30); do + curl -fsS http://127.0.0.1:9000/minio/health/live >/dev/null 2>&1 && break + sleep 2 + [ "$i" = 30 ] && { give_up "minio never became healthy"; return 1; } + done + export RCLONE_CONFIG_EU_TYPE=s3 RCLONE_CONFIG_EU_PROVIDER=Minio \ + RCLONE_CONFIG_EU_ACCESS_KEY_ID=euprobe \ + RCLONE_CONFIG_EU_SECRET_ACCESS_KEY=euprobe123 \ + RCLONE_CONFIG_EU_ENDPOINT=http://127.0.0.1:9000 + rclone mkdir eu:euprobe || give_up "rclone mkdir failed" || return 1 + rclone mount --vfs-cache-mode full --daemon eu:euprobe "$MNT" \ + || give_up "rclone mount failed" || return 1 + defer_shell "fusermount3 -u '$MNT' || fusermount -u '$MNT' || sudo umount -l '$MNT'" + sleep 3 +} + +# Lustre has no single-node story on a stock runner: the client is an +# out-of-tree DKMS module and the server needs either a patched kernel +# (ldiskfs) or ZFS. This probe answers only the first half -- do the +# client modules build against the runner's kernel at all? -- because +# that is the gate everything else is behind. +setup_lustre_client() { + local codename repo + # shellcheck disable=SC1091 # /etc/os-release is provided by the OS + codename="$(. /etc/os-release && echo "$VERSION_ID" | tr -d .)" + repo="https://downloads.whamcloud.com/public/lustre/latest-release/ubuntu${codename}/client" + echo "I: checking $repo" + if ! curl -fsI "$repo/" >/dev/null 2>&1; then + give_up "no Whamcloud client repo for ubuntu${codename} (kernel $(uname -r))" + return 1 + fi + echo "I: module packages published in that repo:" + curl -fsS "$repo/Packages" 2>/dev/null \ + | sed -n 's/^Package: \(lustre-client-modules.*\)/ \1/p' | sort -u | head -20 + apt_install dkms "linux-headers-$(uname -r)" || return 1 + echo "deb [trusted=yes] $repo/ ./" | sudo tee /etc/apt/sources.list.d/lustre.list >/dev/null + defer_shell "sudo rm -f /etc/apt/sources.list.d/lustre.list" + apt_install lustre-client-modules-dkms lustre-client-utils \ + || give_up "lustre client DKMS build failed on kernel $(uname -r)" || return 1 + sudo modprobe lustre || give_up "lustre module built but will not load" || return 1 + NOTE="client modules load; no server (needs patched kernel or ZFS OSD)" + # There is no filesystem to mount: a client without a server is as + # far as a single runner goes. Report what we learned and stop. + return 2 +} + +# OpenAFS is the other out-of-tree client in this space, and unlike +# Lustre it is packaged by Debian/Ubuntu themselves, so the module is +# built by DKMS against whatever kernel the runner has. +setup_openafs() { + apt_install dkms "linux-headers-$(uname -r)" openafs-modules-dkms openafs-client \ + || give_up "openafs DKMS build failed on kernel $(uname -r)" || return 1 + sudo modprobe openafs || give_up "openafs module built but will not load" || return 1 + NOTE="client module builds and loads; a cell (server) is a separate problem" + return 2 +} + +# Filesystems that categorically need a VM (or another machine) rather +# than a runner: report what the runner offers a VM-based backend. +setup_vm_only() { + echo "I: /dev/kvm: $(test -e /dev/kvm && echo present || echo absent)" + echo "I: qemu-system-x86_64: $(command -v qemu-system-x86_64 || echo absent)" + echo "I: nested virt would be needed for: 9p, virtiofs, Lustre server, GPFS" + NOTE="informational only" + return 2 +} + +# ---------------------------------------------------------------- driver + +echo "=== probing candidate: $CANDIDATE" +# shellcheck disable=SC1091 # /etc/os-release is provided by the OS +echo "=== runner: $(. /etc/os-release && echo "$PRETTY_NAME"), kernel $(uname -r)" + +start=$(date +%s) +case "$CANDIDATE" in + gfs2) setup_gfs2 ;; + ocfs2) setup_ocfs2 ;; + f2fs) setup_f2fs ;; + exfat) setup_exfat ;; + ntfs3) setup_ntfs3 ;; + udf) setup_udf ;; + nilfs2) setup_nilfs2 ;; + bcachefs) setup_bcachefs ;; + zfs) setup_zfs ;; + overlay) setup_overlay ;; + glusterfs) setup_glusterfs ;; + cephfs) setup_cephfs ;; + cifs) setup_cifs ;; + sshfs) setup_sshfs ;; + gocryptfs) setup_gocryptfs ;; + encfs) setup_encfs ;; + ecryptfs) setup_ecryptfs ;; + s3-rclone) setup_s3_rclone ;; + lustre-client) setup_lustre_client ;; + openafs) setup_openafs ;; + vm-only) setup_vm_only ;; + *) echo "unknown candidate: $CANDIDATE" >&2; usage >&2; exit 2 ;; +esac +rc=$? +elapsed=$(( $(date +%s) - start )) + +caps="" +case "$rc" in + 0) + echo "=== mounted, collecting capabilities" + findmnt -o SOURCE,TARGET,FSTYPE,OPTIONS "$MNT" || true + caps="$("$HERE/fs-capabilities.sh" "$MNT" 2>&1)" + echo "$caps" + verdict="BOOTSTRAPPED" + ;; + 2) verdict="PARTIAL" ;; + *) verdict="NOT BOOTSTRAPPABLE" ;; +esac + +echo "=== $CANDIDATE: $verdict (${elapsed}s) ${NOTE:+-- $NOTE}" + +if [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then + { + echo "### \`$CANDIDATE\`: $verdict (${elapsed}s)" + [ -n "$NOTE" ] && echo "$NOTE" + if [ -n "$caps" ]; then + echo + echo '```' + echo "$caps" + echo '```' + fi + } >> "$GITHUB_STEP_SUMMARY" +fi + +# One machine-readable line, last, so `tail` of a CI log carries the +# entire finding without scrolling back through apt output. +echo "SUMMARY|$CANDIDATE|$verdict|${elapsed}s|${NOTE:--}|$(echo "$caps" \ + | grep -E '^[a-z0-9-]+=' | tr '\n' ' ')" + +# Every verdict is a finding, including "NOT BOOTSTRAPPABLE" -- that is +# the answer to the question this script asks, not a broken build. So we +# exit 0 once a verdict has been recorded, and let the roll-up table say +# what happened. +# +# The alternative was tried and is worse: with lustre, openafs, cephfs +# and ecryptfs permanently unbootstrappable on a hosted runner, exiting +# non-zero pins four red checks to every pull request that touches +# bin/, and "are this PR's checks green?" stops having an answer. +# +# Genuine script errors (no candidate given, unknown candidate, an +# unusable probe dir) still exit non-zero -- they return before here. +exit 0 diff --git a/bin/ci/render-probe-table.sh b/bin/ci/render-probe-table.sh new file mode 100755 index 0000000..4dca4e7 --- /dev/null +++ b/bin/ci/render-probe-table.sh @@ -0,0 +1,53 @@ +#!/bin/bash +# SPDX-FileCopyrightText: 2026 Yaroslav Halchenko +# SPDX-License-Identifier: MIT +# +# Generated with Claude Code +# +# Render the probe run's per-candidate SUMMARY lines as one Markdown +# table plus the raw capability profiles, so answering "which candidates +# came up?" does not mean opening two dozen job logs. +# +# Reads the artifacts that .github/workflows/probe-filesystems.yaml +# downloads, one summary.txt per candidate, each holding a single line: +# +# SUMMARY||||