Skip to content

Latest commit

 

History

6 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

metal-inference

An agent skill (Claude Code, and any harness that reads SKILL.md) for running local LLMs on Apple Silicon: whether a model/quant/context fits unified memory, which engine and quant to use, launch commands, and triage of slow or OOM-ing models.

Install

git clone https://github.com/PyModel/metal-inference.git ~/.claude/skills/metal-inference

Scripts

  • scripts/envelope.sh — read-only snapshot: chip, RAM, wired limit, memory in use, GPU-held memory and its processes, swap, installed engines, busy ports. Safe while a server holds the GPU.
  • scripts/fit.py — memory budget and verdict against both RAM and the GPU wired cap; measures other processes itself. fit.py --help; fit.py --self-test.
scripts/fit.py --gguf model.gguf --ctx 32768 --layers 48 --kv-heads 8 --head-dim 128

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages