Skip to content

Latest commit

 

History

History
57 lines (39 loc) · 4.18 KB

File metadata and controls

57 lines (39 loc) · 4.18 KB

Segmentation Model Testing

There is no separate test mode

main.py has no --test or --eval flag. The test split is evaluated inside BaseTrainer.train(), once, at the end of the final epoch (base/base_trainer.py):

if epoch == self.epochs:
    self.eval_epoch(epoch, 'test')

Consequences:

  • The test set is evaluated with the model weights from the last epoch. model_best.pth is not reloaded for testing, so the test metrics describe model_last.pth, not the best validation checkpoint.
  • To get test metrics you must run training to completion. The command is the one in training.md; a run finishes when it reaches --epochs.
  • Resuming a finished run does not run the test. --resume restarts from checkpoint['epoch'] + 1. If the stored epoch already equals --epochs, the training loop has no iterations and eval_epoch(..., 'test') is never called. If you pass a larger --epochs, training continues and the test runs at the new final epoch.
  • An interrupted run that is resumed with the same --epochs evaluates the test set normally when it reaches the last epoch.

Which metrics are computed on the test set is controlled by the "test": true flag of each entry in config.metrics (see config/README.md). For volumetric trainers (Trainer_3D, Trainer_3DText), test inference is patch-based: each full volume is covered with a TorchIO GridSampler (dataset.patch_size, dataset.grid_overlap), and the patch predictions are recombined with a GridAggregator (Hann overlap).

If W&B logging is enabled (--wandb), the test metrics are also logged under the test/ prefix.

Output metric CSVs

MetricsManager.save_to_csv writes one CSV per phase to --save_path:

File Written
train_metrics.csv After every training epoch.
val_metrics.csv After every validation run (every --val_every epochs, only with --validation).
test_metrics.csv Once, after the test evaluation at the final epoch.

Each file has one row per evaluated epoch, with no per-case rows. Each value is the mean over all batches or volumes of that epoch, all-reduced across DDP ranks. The columns are:

Column Meaning
epoch Epoch number.
<loss name> Loss value (e.g. DiceFocalLoss). Only in the train and val files, because the loss is not added to the test metrics.
<key>_<class name> Per-class value of metric <key> (e.g. DSC_Edema). Class names come from config.classes, or class_N if a class has no name. If the metric excludes the background, the per-class columns start at class 1.
<key>_mean Mean over the per-class values.
<key>_<region> Value on each aggregated region from config.aggregated_regions (e.g. DSC_ET, DSC_TC, DSC_WT). Present only if the config defines aggregated_regions.
<key>_aggregated_mean Mean over the aggregated regions.

<key> is the key field of a metric entry in config.metrics (e.g. DSC, HD95). On --resume, the existing CSVs are loaded back so that later rows are appended.

Using the results for statistical comparison

statistical_test.py compares several models with a paired test (see statistical_test.md). You pass one CSV per model and choose a column with --metric-column (e.g. DSC_aggregated_mean or DSC_mean, the same <key>_mean / <key>_aggregated_mean names as above).

statistical_test.py needs per-case CSVs. Each file must have one row per test subject, identified by --id-column (default subject_id). An optional AVERAGE row is dropped automatically. The test_metrics.csv written by the trainers holds only one aggregated row per epoch and has no subject_id column, so you cannot pass it to statistical_test.py directly. The per-case tables must be produced separately. The model name used in the comparison comes from each CSV's file name, so give every model's file a distinct name.

Related Documentation