Skip to content

Support for YOLO-style / modern object detection models (2D and 3D) #9111

Description

@vikashg

Is your feature request related to a problem? Please describe.

MONAI currently ships a single object-detection architecture — RetinaNet
under monai/apps/detection/ (retinanet_detector.py, retinanet_network.py),
with supporting anchor utilities, COCO-style mAP metrics, and box transforms. This
is a solid anchor-based, two-stage-style foundation and works in both 2D and 3D.

However, the detector zoo stops there. Users who want fast single-stage or
anchor-free detectors most notably the YOLO family — have no native path.
The existing guidance (see #903) is essentially "use MONAI transforms inside your
own YOLO pipeline," which leaves the detector itself, training loop, box-format
handling, and metrics outside MONAI's guarantees. #292 raised YOLO/COCO/Pascal-VOC
box-format support in transforms years ago but there is no dedicated tracking
issue for native modern detectors, and #8519 lists surgical instrument
localization/detection as a target without naming an architecture.

Describe the solution you'd like

Expand monai/apps/detection beyond RetinaNet to include modern detectors, with
YOLO as the flagship because of its strong fit for 2D medical / endoscopy /
surgical-tool and microscopy use cases (real-time inference, anchor-free variants,
mature ecosystem). Concretely:

  1. A YOLO-style detector network (e.g. an anchor-free YOLO head) integrated into
    the existing DetectorNetwork / detector API so it reuses MONAI's box
    transforms, anchor/anchor-free utilities, ATSS-style matching, and mAP metrics.
  2. Native handling of the YOLO box format (normalized cx, cy, w, h) in the
    detection transforms, alongside the existing corner/CCWH conventions (follow-up
    to 3D transform for detection task #292).
  3. A tutorial mirroring the existing RetinaNet LUNA16 / detection tutorial so the
    new detector is a drop-in alternative.
  4. (Optional / stretch) Room in the API for other modern detectors such as
    DETR-family or FCOS, so this is an extensible "detector zoo" rather than a
    one-off.

Note: interest in YOLO-inspired 3D detection

MONAI's biggest differentiator over general-purpose CV libraries is first-class
3D support — RetinaNet here already runs on volumetric data. If there is
community/maintainer interest, a YOLO-inspired 3D detector (a single-stage /
anchor-free volumetric detection head operating on 3D feature maps, predicting
6-DoF axis-aligned 3D boxes) would be a genuinely novel and high-value addition:
few libraries offer a fast single-stage 3D detector, and use cases like nodule /
lesion / landmark detection in CT and MR volumes would benefit directly. I'd be
happy to help scope and prototype this if the maintainers see value — please
comment if there's appetite for the 3D direction specifically.

Describe alternatives you've considered

  • Continuing to use RetinaNet only (works, but no fast single-stage/anchor-free
    option, and no YOLO ecosystem interop).
  • Running an external YOLO (ultralytics etc.) alongside MONAI purely for transforms
    (the current can use yolo in monai ? #903 answer) — loses MONAI's metric/box/3D guarantees and
    reproducibility.

Additional context

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions