Skip to content

Auto model selection picks models too weak for Linux kernel patch-series work #4917

Description

@rppt

Across many sessions on Linux kernel patch series, Auto mode selects models in
the gpt-5.6-sol class that consistently fail at the core skill this work
requires: respecting commit boundaries and scope. This is a recurring pattern,
not a one-off. Stronger models handle the same tasks reliably.

Typical failure modes, all occurring with explicit standing instructions in
context ("fold changes into the relevant commit", "keep history bisectable",
"no functional change intended"):

  1. Scope creep inside code-motion commits. A commit meant to relocate code also
    introduces new abstractions or refactors. Correcting this usually takes
    several rounds, because the model re-expands scope after being told to
    narrow it.

  2. Silent deletion or open-coding of helpers. Existing static helpers get
    inlined away during "no functional change" commits - unrequested and
    unmentioned, discoverable only by manual diff review.

  3. Incomplete renames. A requested rename is applied to the obvious symbols but
    leaves related identifiers behind, often in other architectures or headers.

  4. Architecture-correctness misses. Code carrying one architecture's semantics
    is lifted into generic code and applied to others, without recognizing the
    semantics were not universal.

  5. Misread API shape. Instructions about where code should live (per-architecture
    copies vs. a shared implementation with wrappers) are inverted.

  6. Changing established semantics to fit the model's design. Long-standing
    function behavior is altered to make a new abstraction fit, instead of
    preserving behavior.

The common thread: these models optimize for a plausible-looking final tree and
treat the commit series as an implementation detail. For kernel work the series
is the deliverable.

Requests:

  • Weight Auto model selection toward stronger models for kernel/systems C work,
    especially sessions involving git rebase, commit splitting, or
    cross-architecture code.
  • Surface which model Auto selected, so a mismatch is visible before hours are
    spent.
  • Treat "fold into the relevant commit" and "no functional change intended" as
    hard constraints, not stylistic hints.

Kernel maintainers evaluating Copilot CLI judge it on whether a series is clean,
bisectable and correctly scoped. A model that quietly reshapes commits fails
that bar regardless of how good the final tree looks.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions