main.py has no --test or --eval flag. The test split is evaluated inside BaseTrainer.train(), once, at the end of the final epoch (base/base_trainer.py):
if epoch == self.epochs:
self.eval_epoch(epoch, 'test')Consequences:
- The test set is evaluated with the model weights from the last epoch.
model_best.pthis not reloaded for testing, so the test metrics describemodel_last.pth, not the best validation checkpoint. - To get test metrics you must run training to completion. The command is the one in training.md; a run finishes when it reaches
--epochs. - Resuming a finished run does not run the test.
--resumerestarts fromcheckpoint['epoch'] + 1. If the stored epoch already equals--epochs, the training loop has no iterations andeval_epoch(..., 'test')is never called. If you pass a larger--epochs, training continues and the test runs at the new final epoch. - An interrupted run that is resumed with the same
--epochsevaluates the test set normally when it reaches the last epoch.
Which metrics are computed on the test set is controlled by the "test": true flag of each entry in config.metrics (see config/README.md). For volumetric trainers (Trainer_3D, Trainer_3DText), test inference is patch-based: each full volume is covered with a TorchIO GridSampler (dataset.patch_size, dataset.grid_overlap), and the patch predictions are recombined with a GridAggregator (Hann overlap).
If W&B logging is enabled (--wandb), the test metrics are also logged under the test/ prefix.
MetricsManager.save_to_csv writes one CSV per phase to --save_path:
| File | Written |
|---|---|
train_metrics.csv |
After every training epoch. |
val_metrics.csv |
After every validation run (every --val_every epochs, only with --validation). |
test_metrics.csv |
Once, after the test evaluation at the final epoch. |
Each file has one row per evaluated epoch, with no per-case rows. Each value is the mean over all batches or volumes of that epoch, all-reduced across DDP ranks. The columns are:
| Column | Meaning |
|---|---|
epoch |
Epoch number. |
<loss name> |
Loss value (e.g. DiceFocalLoss). Only in the train and val files, because the loss is not added to the test metrics. |
<key>_<class name> |
Per-class value of metric <key> (e.g. DSC_Edema). Class names come from config.classes, or class_N if a class has no name. If the metric excludes the background, the per-class columns start at class 1. |
<key>_mean |
Mean over the per-class values. |
<key>_<region> |
Value on each aggregated region from config.aggregated_regions (e.g. DSC_ET, DSC_TC, DSC_WT). Present only if the config defines aggregated_regions. |
<key>_aggregated_mean |
Mean over the aggregated regions. |
<key> is the key field of a metric entry in config.metrics (e.g. DSC, HD95). On --resume, the existing CSVs are loaded back so that later rows are appended.
statistical_test.py compares several models with a paired test (see statistical_test.md). You pass one CSV per model and choose a column with --metric-column (e.g. DSC_aggregated_mean or DSC_mean, the same <key>_mean / <key>_aggregated_mean names as above).
statistical_test.py needs per-case CSVs. Each file must have one row per test subject, identified by --id-column (default subject_id). An optional AVERAGE row is dropped automatically. The test_metrics.csv written by the trainers holds only one aggregated row per epoch and has no subject_id column, so you cannot pass it to statistical_test.py directly. The per-case tables must be produced separately. The model name used in the comparison comes from each CSV's file name, so give every model's file a distinct name.