Skip to content

About

Run large AI models on your own PC: Qwen3.8-Flash-Next, GLM 5.3 Flash, Qwen3.6, Ornith 1.5 and Gemma 4. One-click install for Windows and Linux, OpenAI/Anthropic API on localhost, image input. Built on Strata.

Topics

Resources

Security policy

Stars

66 stars

Watchers

3 watching

Forks

Latest commit

 

History

1,992 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

PolyStrata

English · 简体中文 · 日本語 · Deutsch · Français · Español · Português

Run large AI models on your own PC
Qwen3.8-Flash-Next, GLM 5.3 Flash, Qwen3.6, Ornith 1.5 and Gemma 4 · free and open source

Tokens per second while writing an answer: 84 to 87 for Qwen3.6-35B-A3B on a laptop with an 8 GB card, 385 to 393 on an RTX 5090

Claude Code in a terminal beside PolyStrata's API Monitor
Claude Code on Qwen3.6-35B-A3B UD-Q3_K_M, on the laptop with the 8 GB card, in real time: it adds a --json option to a small program, writes a test for it and runs the tests. The server had read Claude Code's system prompt when it started ("warm_start": true).

A voxel pagoda garden the model wrote, running in the browser
A voxel pagoda garden, 1 shot prompt running on an RTX 5070 with Strata (Qwen3.8-Flash-Next IQ3_S, 128K context) · full video (49 s)

PolyStrata runs large AI models on a normal PC, models that usually need a server. They chat, write code and work with your apps and coding agents. Nothing leaves your PC.

It runs five models, on two engines:

  • Qwen3.8-Flash-Next on an NVIDIA or AMD graphics card with 12 GB or more, on Windows or Linux. It also reads pictures. Its engine is the Strata project's, unchanged.

The other four run on an engine written for PolyStrata, with one NVIDIA card: on Linux, and the three smaller ones on Windows too. They are experimental: they have run on one PC with a 32 GB card, and the three smaller ones on a laptop with an 8 GB card as well.

  • GLM 5.3 Flash, the largest: a file of 79 to 114 GB, most of which stays in RAM.
  • Qwen3.6-35B-A3B, a file of 13 to 23 GB. It also reads pictures.
  • Ornith 1.5, which has Qwen3.6-35B-A3B's architecture and other weights: a file of 14 to 24 GB. It also reads pictures.
  • Gemma 4 26B-A4B, a file of 11 to 17 GB. It also reads pictures.

The server can change between installed models while it runs (how).

PolyStrata is built on Strata by Niko1221. Why it is a separate repo, AMD and smaller cards, fine-tunes and other questions: FAQ.

How fast is it?

A token is about ¾ of a word.

  • Writes answers: how fast the reply appears in a short chat. 60 tokens per second is faster than you can read.
  • Reads your prompt: how fast it takes in what you send (a document, code or a chat history).

Qwen3.8-Flash-Next

Measured by the Strata project on two ordinary gaming PCs, with a 32K-token prompt.

NVIDIA: RTX 5070 (12 GB), Ryzen 5 7600, 64 GB RAMAMD: RX 9070 XT (16 GB), Ryzen 9 3900X, 47 GB RAM
Size Writes answers Reads your prompt
Q2_0 94 tokens/s 2,650 tokens/s
IQ2_XS 79 tokens/s 2,090 tokens/s
IQ3_XXS 62 tokens/s 1,750 tokens/s
IQ3_S 53 tokens/s 1,620 tokens/s
Coder 55 tokens/s 2,180 tokens/s
Size Writes answers Reads your prompt
Q2_0 60 tokens/s 1,160 tokens/s
IQ2_XS 52 tokens/s 1,110 tokens/s
Coder 44 tokens/s 1,420 tokens/s

NVIDIA: Q2_0 with engine 0.1.36, the other rows with 0.1.26 (4K answers, 32K prompts). The full tables are in DETAILS.md. A card with more VRAM is faster: an RTX 3090 (24 GB) should write about 100-140 tokens per second. Long chats and other cards: speed of each size, community results.

GLM 5.3 Flash

Measured for PolyStrata on one workstation, with a short chat and a 29K-token prompt.

NVIDIA: RTX 5090 (32 GB), Xeon Platinum 8470Q (25 vCPU), 120 GB RAM
Size Writes answers Reads your prompt
3.0BIT 59-66 tokens/s 820 tokens/s
UD-IQ2_XXS 47-50 tokens/s 1,040-1,050 tokens/s
REAP50-Q3_K_M 59-66 tokens/s 1,300-1,310 tokens/s

GLM 5.3 Flash is a much larger model than the others, and about a fifth of its experts fit in this card's 32 GB, so it writes more slowly. What was measured and how: docs/GLM.md.

Qwen3.6-35B-A3B

Measured for PolyStrata on two PCs, with a short chat and a prompt of 30,000 to 32,000 tokens: on the left a laptop whose card has 8 GB, where the processor runs the experts the card has no room for, and on the right a workstation whose card holds the whole model.

NVIDIA: RTX 4070 Laptop (8 GB), Core i7-14700HX, 32 GB RAMNVIDIA: RTX 5090 (32 GB), Xeon Platinum 8470Q (25 vCPU), 120 GB RAM
Size Writes answers Reads your prompt
UD-Q3_K_M 84-87 tokens/s 1,500-1,620 tokens/s
UD-Q2_K_XL 110-111 tokens/s 1,500-1,510 tokens/s
UD-Q4_K_M 79-86 tokens/s 1,560-1,570 tokens/s
Size Writes answers Reads your prompt
UD-Q3_K_M 385-393 tokens/s 6,350-6,550 tokens/s
UD-Q2_K_XL 410-438 tokens/s 5,970-6,440 tokens/s
UD-Q4_K_M 328-393 tokens/s 6,230-6,580 tokens/s

A prompt of 10,000 tokens is read at 1,620 tokens/s on the laptop and at 6,100 on the workstation. The workstation's card held to 8 GB, a stand-in from before the laptop was measured, is in docs/BACKENDS.md. The laptop was measured with --tail-skip off, which the installer now turns on for Qwen3.6 (docs/BACKENDS.md).

Ornith 1.5

Measured the same way, on the same two PCs.

NVIDIA: RTX 4070 Laptop (8 GB), Core i7-14700HX, 32 GB RAMNVIDIA: RTX 5090 (32 GB), Xeon Platinum 8470Q (25 vCPU), 120 GB RAM
Size Writes answers Reads your prompt
APEX-I-COMPACT 72-76 tokens/s 1,540 tokens/s
APEX-I-MINI 77-79 tokens/s 1,470 tokens/s
APEX-I-QUALITY 59-63 tokens/s 940-1,510 tokens/s
Size Writes answers Reads your prompt
APEX-I-COMPACT 303-315 tokens/s 6,080-6,360 tokens/s
APEX-I-MINI 289-300 tokens/s 5,940-6,170 tokens/s
APEX-I-QUALITY 289-298 tokens/s 6,040-6,390 tokens/s

A prompt of 10,000 tokens is read at 1,530 tokens/s on the laptop and at 5,600 on the workstation.

Gemma 4 26B-A4B

Measured the same way, on the same two PCs.

NVIDIA: RTX 4070 Laptop (8 GB), Core i7-14700HX, 32 GB RAMNVIDIA: RTX 5090 (32 GB), Xeon Platinum 8470Q (25 vCPU), 120 GB RAM
Size Writes answers Reads your prompt
UD-Q3_K_M 85 tokens/s 2,100 tokens/s
UD-Q2_K_XL 94-95 tokens/s 1,920-2,000 tokens/s
UD-Q4_K_M 74-75 tokens/s 2,010 tokens/s
Size Writes answers Reads your prompt
UD-Q3_K_M 326-349 tokens/s 7,880-8,230 tokens/s
UD-Q2_K_XL 316-331 tokens/s 7,120-7,860 tokens/s
UD-Q4_K_M 316-338 tokens/s 7,040-8,050 tokens/s

A prompt of 10,000 tokens is read at 2,080 tokens/s on the laptop and at 6,800 on the workstation.

The four models of PolyStrata's own engine were measured with its benchmark, which you can run on your PC: python polystrata.py benchmark --model <name> --prompt-tokens 24000 (what it sends and prints). What others measured on their own cards, and how to compare with llama.cpp on yours: community results.

What you need

Qwen3.8-Flash-Next

Graphics card NVIDIA GeForce RTX 20, 30, 40 or 50 series, or AMD Radeon RX 7900 XT / XTX, RX 7800 XT / 7700 XT, RX 9060 XT, RX 9070 / 9070 XT, Radeon AI PRO R9700 or RX 6800 / 6900 series. It needs 12 GB of VRAM or more.
RAM 32 GB or more. Your RAM decides which size fits. 64 GB runs every size.
Disk About 80 GB free. Use an SSD if you can: the first start is much faster.
System Windows 10 / 11 or Linux, and a current graphics driver from NVIDIA or AMD.

The installer sets up everything else. Two or three cards can share the model (multi-GPU).

Experimental, written and tested by community members on their own machines:

  • Older graphics cards (Tesla P40 / V100, GTX 10, Radeon VII / MI50, RX 6700 XT, RX 5500 XT): Older GPUs.
  • Intel Arc, built from source on Linux: Intel Arc.
  • AMD Ryzen AI Max (Strix Halo), built from source on Linux: Strix Halo.
  • Older processors without AVX2: they work, but slowly. Older CPUs.

The full list: docs/INSTALL.md.

GLM 5.3 Flash

Graphics card One NVIDIA GeForce RTX 20 series card or newer. It has run on 32 GB of VRAM only. About 12 GB are taken before any expert, so a 12 GB card cannot start it.
RAM Beside a 32 GB card, 96 GB or more for the whole model, 82 GB for its 2-bit size, 60 GB for the size with half the experts. A smaller card needs more RAM (how much). It has run with 120 GB.
Disk 114 GB, 102 GB or 79 GB for the model file, on an SSD.
System Linux, and a current NVIDIA driver.

The installer sets up everything else. It compiles the engine on your PC and offers to install what that needs (g++ and the CUDA toolkit). Not yet: Windows, AMD cards, several cards, pictures.

The full list: docs/GLM.md.

Qwen3.6-35B-A3B

Graphics card One NVIDIA GeForce RTX 20 series card or newer with 8 GB of VRAM or more. About 5 GB are taken before any expert. It has run on a 32 GB card, on that card held to 8 GB, and on a laptop's 8 GB card.
RAM About 17, 21 or 27 GB beside an 8 GB card for the three sizes, the installer's estimate. A card with 24 GB holds the two smaller sizes whole. All three sizes have run on a laptop with 32 GB.
Disk 13, 17 or 23 GB for the model file, 1 GB more to read pictures.
System Linux or Windows 10 / 11, and a current NVIDIA driver.

The installer sets up everything else, as it does for GLM 5.3 Flash. On Windows it downloads the engine ready-made from the latest release. On a processor without AVX2, or in a copy whose engine changed after that release, it compiles the engine instead with Visual Studio's 2022 compiler and the CUDA toolkit, and offers to install them. Not yet: AMD cards, several cards.

Ornith 1.5

Graphics card One NVIDIA GeForce RTX 20 series card or newer with 8 GB of VRAM or more. About 4 GB are taken before any expert. It has run on a 32 GB card, on that card held to 8 GB, and on a laptop's 8 GB card.
RAM About 18, 22 or 28 GB beside an 8 GB card for the three sizes, the installer's estimate. A card with 24 GB holds the two smaller sizes whole. All three sizes have run on a laptop with 32 GB.
Disk 14, 17 or 24 GB for the model file, 1 GB more to read pictures.
System Linux or Windows 10 / 11, and a current NVIDIA driver.

The installer sets up everything else, as it does for GLM 5.3 Flash. On Windows it downloads the engine ready-made from the latest release. On a processor without AVX2, or in a copy whose engine changed after that release, it compiles the engine instead with Visual Studio's 2022 compiler and the CUDA toolkit, and offers to install them. Not yet: AMD cards, several cards.

Gemma 4 26B-A4B

Graphics card One NVIDIA GeForce RTX 20 series card or newer with 8 GB of VRAM or more. About 6 GB are taken before any expert. It has run on a 32 GB card, on that card held to 8 GB, and on a laptop's 8 GB card.
RAM About 15, 18 or 21 GB beside an 8 GB card for the three sizes, the installer's estimate. By the same estimate a card with 16 GB holds the two smaller sizes whole. All three sizes have run on a laptop with 32 GB.
Disk 11, 13 or 17 GB for the model file and its draft model, 1 GB more to read pictures.
System Linux or Windows 10 / 11, and a current NVIDIA driver.

The installer sets up everything else, as it does for GLM 5.3 Flash. On Windows it downloads the engine ready-made from the latest release. On a processor without AVX2, or in a copy whose engine changed after that release, it compiles the engine instead with Visual Studio's 2022 compiler and the CUDA toolkit, and offers to install them. Not yet: AMD cards, several cards.

What each of the three needs and how it was checked: docs/BACKENDS.md.

Install

Let your AI set it up

Do you use an AI coding assistant (Claude Code, Cursor, Codex, GitHub Copilot, ...)? Paste this into it:

Set up PolyStrata on this PC for me: https://github.com/VecSzn/PolyStrata - follow docs/AI_SETUP.md in that repository.

It checks your graphics card, RAM and disk and picks the model that fits. Then it installs and starts it and tells you how to connect your apps. AI tools can also install, start and stop PolyStrata through its MCP server.

Or do it yourself

Download PolyStrata and unzip it (or git clone it). Windows: double-click START-HERE.bat. Linux: run ./setup.sh in the PolyStrata folder.

The installer finds your card and sets up the right engine for it. It asks you a few questions:

  • which model and which size,
  • how much context (how much text the model keeps in mind),
  • whether it should read pictures (GLM 5.3 Flash reads none).

Press Enter each time for the recommended answer. Then it downloads the model (about 70 GB for Qwen3.8-Flash-Next, 79 to 114 GB for GLM 5.3 Flash, 11 to 24 GB for the other three) and starts it. If the download stops, run it again: it continues where it left off. Your browser opens the PolyStrata app at http://127.0.0.1:8080.

All five are in the installer's list of models. ./setup.sh --family glm, qwen36, ornith or gemma4 goes straight to one (what it asks and writes).

While the model starts, your PC can be slow or stop responding for 1-3 minutes (longest the first time). PolyStrata loads 35-55 GB into your RAM and locks part of it for the graphics card. This is normal. Wait, and don't close the window. The window shows what PolyStrata is doing.

Next time, run START-HERE.bat (or ./setup.sh) again. It starts right away and downloads nothing twice. Close its window to stop the model. UPDATE.bat (./update.sh) updates PolyStrata without starting it. Updating, Docker, several cards, where the files go and every option: docs/INSTALL.md.

Which model should I pick?

Qwen3.8-Flash-Next

The installer recommends one for your RAM. The same model comes in several sizes, compressed more or less. Smaller sizes are faster. Larger sizes are a bit smarter.

Your RAM Take Why
32 GB Coder it fits 32 GB, and it is made for code (with a 24 GB card, Q2_0 and IQ2_XS run too)
48 GB IQ2_XS (or Q2_0, the fastest) the larger sizes do not fit
64 GB IQ2_XS (recommended), or IQ3_XXS / IQ3_S every size fits; IQ3_S is the best and the slowest
96 GB or more IQ3_S, or Unsloth's UD-IQ4_XS (~4-bit) room for the largest sizes with everything else open
  • Coder: a coding version with half of the experts removed. It reaches 91% of the full model's SWE-bench Verified score (measured by its authors) and fits 32 GB of RAM. It is weaker outside code, including Chinese and other CJK text (#438). For those, take Q2_0, IQ2_XS or IQ3_S, which keep every expert.
  • Swift 1.5: a fine-tune that thinks for a much shorter time before it answers. You get the answer sooner, at about the same quality.
  • Unsloth UD-IQ4_XS: Unsloth's ~4-bit version, between IQ3_S and UD-Q4_K_XL in quality. A 94 GB download. With less than ~80 GB of RAM, PolyStrata reads part of it from the SSD while it answers, so it is slower there (an NVMe SSD helps).
  • Unsloth UD-Q4_K_XL (experimental): the closest to the full model. But PolyStrata reads most of it from the SSD while it answers, so it writes only 7-8.5 tokens/s on a 64 GB PC.
  • OrcaRouter's Uncensored IQ3_XXS: you set it up by hand. It is not in the installer's menu.

Sizes, downloads and what fits where: docs/MODELS.md. To add another model later, run SETUP.bat (Linux: ./setup.sh --setup).

GLM 5.3 Flash

Three sizes so far. The installer takes the first one your RAM is enough for.

Size What it is Download RAM beside a 32 GB card
3.0BIT the whole model, at about 3 bits (neuralll) 114 GB 96 GB
UD-IQ2_XXS the whole model, at about 2 bits (Unsloth): it reads a prompt faster and writes more slowly 102 GB 82 GB
REAP50-Q3_K_M half of the routed experts removed (REAP50 by patrickbdevaney) 79 GB 60 GB

Your graphics card decides how much RAM it needs, because the experts the card does not hold have to stay in RAM.

Your graphics card 3.0BIT UD-IQ2_XXS REAP50-Q3_K_M
32 GB 96 GB 82 GB 60 GB the PC it has run on (120 GB of RAM)
24 GB about 104 GB about 90 GB about 68 GB not run yet
16 GB about 113 GB about 99 GB about 77 GB not run yet
12 GB or less it does not start

The RAM numbers are the installer's estimate from the one PC it has run on. With less RAM it still starts, but it reads experts from the SSD while it answers and gets much slower. The installer says so and asks first.

Other quantizations of GLM 5.3 Flash are not supported yet: what is missing.

Qwen3.6-35B-A3B

Three sizes so far. The installer takes UD-Q3_K_M, or UD-Q2_K_XL when the PC has too little RAM for it. UD-Q4_K_M is chosen by name: --model UD-Q4_K_M.

Size What it is Download RAM beside an 8 GB card
UD-Q3_K_M the whole model with 3-bit experts (Unsloth) 17 GB about 21 GB
UD-Q2_K_XL the same with 2- and 3-bit experts: the least RAM 13 GB about 17 GB
UD-Q4_K_M the same with 4-bit experts 23 GB about 27 GB

A larger card holds more of the model and needs less RAM: with UD-Q3_K_M about 17 GB beside a 12 GB card, 13 GB beside a 16 GB card, and 10 GB once the card holds all of it. These are the installer's estimates. They have not been checked on such a PC.

Ornith 1.5

Three sizes so far. The installer takes APEX-I-COMPACT, or APEX-I-MINI when the PC has too little RAM for it. APEX-I-QUALITY is chosen by name: --model APEX-I-QUALITY.

Size What it is Download RAM beside an 8 GB card
APEX-I-COMPACT the whole model with 3- and 4-bit experts (APEX by mudler) 17 GB about 22 GB
APEX-I-MINI the same with 2- and 3-bit experts: the least RAM 14 GB about 18 GB
APEX-I-QUALITY the same with 4- to 6-bit experts 24 GB about 28 GB

A larger card holds more of the model and needs less RAM: with APEX-I-COMPACT about 17 GB beside a 12 GB card, 13 GB beside a 16 GB card, and 10 GB once the card holds all of it. These are the installer's estimates. They have not been checked on such a PC.

Gemma 4 26B-A4B

Three sizes so far. The installer takes UD-Q3_K_M, or UD-Q2_K_XL when the PC has too little RAM for it. UD-Q4_K_M is chosen by name: --model UD-Q4_K_M.

Size What it is Download RAM beside an 8 GB card
UD-Q3_K_M the whole model with 3-bit experts (Unsloth), and Google's draft model for it 13 GB about 18 GB
UD-Q2_K_XL the same with 2-bit experts: the least RAM 11 GB about 15 GB
UD-Q4_K_M the same with 4- and 5-bit experts 17 GB about 21 GB

A larger card holds more of the model and needs less RAM: with UD-Q3_K_M about 14 GB beside a 12 GB card, and 10 GB once the card holds all of it. These are the installer's estimates. They have not been checked on such a PC.

Using it

All five models are used the same way: the same app, the same API.

The app's Monitor tab next to a coding agent
The app's Monitor (left) while a coding agent writes the pagoda garden from the video (right). A screenshot from Strata.

  • In the browser: open http://127.0.0.1:8080. It has Chat, a live Monitor of the model and your GPU/CPU/RAM, and About with the settings and addresses.
  • Your apps and coding agents: add an "OpenAI-compatible" provider with the base URL http://127.0.0.1:8080/v1. Any API key and any model name work.
    • Apps that use Anthropic's API: http://127.0.0.1:8080/v1/messages (Claude Code: ANTHROPIC_BASE_URL=http://127.0.0.1:8080).
    • Codex CLI and other apps that use the OpenAI Responses API: /v1/responses (setup).
  • Thinking: choose off, low, medium or high in the chat menu or in your app's "reasoning effort". Off is the fastest. High is best for hard questions.
  • From your phone or another PC: START-HERE.bat --setup --host 0.0.0.0 --api-key <secret>. Always set a key.

Where they differ:

Qwen3.8-Flash-Next GLM 5.3 Flash Qwen3.6-35B-A3B, Ornith 1.5, Gemma 4 26B-A4B
Pictures Say yes to "Images?" in setup. Then click Picture in the chat, or attach pictures in your app. AMD cards read pictures on Linux through the processor; on Windows they can't yet. Not yet. Say yes to "Images?" in setup, then as with Qwen3.8-Flash-Next. On a card with less than 16 GB the picture encoder runs on the processor, and a picture takes a few seconds.
Several requests One at a time by default, the others wait. "parallel": 2 answers several at once (BATCHING.md); on a 12 GB card each answer gets slower. One at a time, the others wait. One at a time, the others wait.
Several conversations The last one is kept, so follow-up messages start in seconds. Two engine arguments keep several in RAM (details). The last one is kept. Setup asks whether to keep several in RAM, so that an agent and its subagents take turns without their prompts being read again (how). As GLM 5.3 Flash.
Long prompts The first message of a chat is read in full, about 1 minute per 30,000 tokens. The first message is read in full, about 40 seconds per 30,000 tokens. The first message is read in full: 30,000 tokens take 4 to 5 seconds on the workstation and 14 to 20 seconds on the 8 GB laptop.

More: where your chats are stored, the API.

Something went wrong?

  • My PC froze the first time PolyStrata started. This is normal while it loads the model. Wait, and don't close the window. Still frozen after 10 minutes? Restart the PC, close other programs and try again, or pick a smaller size.
  • It stopped while downloading or installing. Run START-HERE.bat (or ./setup.sh) again. It continues where it stopped.
  • It's very slow and the disk light keeps blinking, or it says "the engine stopped unexpectedly". Your PC does not have enough free RAM. Close other programs (browsers use a lot), or pick a smaller size (Q2_0 or IQ2_XS).
  • It says port 8080 is already in use. PolyStrata is already running. Look for its window.

More problems and their fixes: docs/TROUBLESHOOTING.md. Still stuck? Open an issue and attach strata-<model>.log from the PolyStrata folder. Found a security problem? Report it privately: SECURITY.md.

How does it work?

Models like these usually run on servers with hundreds of gigabytes of graphics memory. Your graphics card has 8-32 GB. PolyStrata makes the model fit by sharing the work across your whole PC. Think of a kitchen: the things you use all the time stay on the counter, and the rest waits in the pantry.

Qwen3.8-Flash-Next

The model's 24,576 experts: the busiest on the graphics card, all of them in RAM, a lookup table on the SSD

  • The model is a team of 24,576 small specialists ("experts"). Each word needs only 10 of them.
  • Your graphics card keeps the few thousand experts that are used most often. Your RAM holds all of them, and your processor works on the rest at the same time. Your SSD holds a big lookup table.

A small helper guesses the next words; the big model checks them all at once and keeps the right ones

  • Guess, then check: a small helper guesses the next few words. The big model checks them all at once. You get the same answer, 1.6-1.8x sooner.
  • Long texts are read in big pieces (up to 8,192 tokens at a time), at over 1,000 tokens per second.

The longer explanation: docs/HOW_IT_WORKS.md. Every part and its numbers: the details and the paper.

GLM 5.3 Flash

The model's 12,096 experts: the most used on the graphics card, the rest in RAM, the model's file on the SSD

  • The model is a team of 12,096 small experts, 288 in each of 42 layers. Each word needs only 8 from every layer.
  • Your graphics card keeps the ones used most often, about 2,470 of them on a 32 GB card: the engine counts which ones your work uses and chooses again at the next start. Your RAM holds the rest, and your processor runs them at the same time. Your SSD holds the model's file, which is used where it is, with no second copy; a start takes about 6 seconds once the file is in your RAM.

A block of the model guesses the next words; the model checks them all at once and keeps the right ones

  • Guess, then check: a block of the model made for guessing proposes the next few words. The model checks them all at once. You get the same answer, 1.1-1.2x sooner.
  • Long texts are read in big pieces (1,024 tokens at a time, the experts the card does not hold passing through it), at 780 to 1,250 tokens per second.

Every part and its numbers: docs/GLM.md.

Qwen3.6-35B-A3B

The model's 10,240 experts: the ones an answer keeps using on the graphics card, the rest in RAM, the model's file on the SSD

  • The model is a team of 10,240 small experts, 256 in each of 40 layers. Each word needs only 8 from every layer.
  • Your graphics card keeps as many as fit: about 2,000 of them on an 8 GB card, all of them from 24 GB on. While it writes, the engine brings in the experts the answer keeps choosing, and on the 8 GB card the ones it holds serve about 55% of what an answer uses. Your RAM holds the rest, and your processor runs them at the same time. Your SSD holds the model's file, which is used where it is, with no second copy.

The model's draft block guesses the next words; the model checks them all at once and keeps the right ones

  • Guess, then check: the model's own draft block, which is in its file, proposes the next few words. The model checks them all at once. You get the same answer, 1.2-1.5x sooner.
  • Long texts are read in big pieces (1,024 tokens at a time). When the card does not hold every expert, the others pass through it once for the whole prompt.

Every part and its numbers: docs/BACKENDS.md.

Ornith 1.5

The model's 10,240 experts: the ones an answer keeps using on the graphics card, the rest in RAM, the model's file on the SSD

  • Qwen3.6-35B-A3B's architecture with other weights: a team of 10,240 small experts, 256 in each of 40 layers, and only 8 of every layer for each word.
  • Your graphics card keeps as many as fit: about 2,500 of them on an 8 GB card, all of them from 24 GB on. While it writes, the engine brings in the experts the answer keeps choosing, and on the 8 GB card the ones it holds serve about 60% of what an answer uses. Your RAM holds the rest, and your processor runs them at the same time. Your SSD holds the model's file, which is used where it is, with no second copy.

The model's draft block guesses the next words; the model checks them all at once and keeps the right ones

  • Guess, then check: its draft block is in its file too, and proposes the next few words. It guesses right less often than Qwen3.6-35B-A3B's, so you get the same answer 1.1-1.2x sooner.
  • Long texts are read in big pieces (1,024 tokens at a time). When the card does not hold every expert, the others pass through it once for the whole prompt.

Every part and its numbers: docs/BACKENDS.md.

Gemma 4 26B-A4B

The model's 3,840 experts: the ones an answer keeps using on the graphics card, the rest in RAM, the model's file on the SSD

  • The model is a team of 3,840 small experts, 128 in each of 30 layers. Each word needs only 8 from every layer, beside a part of the layer that every word uses.
  • Your graphics card keeps as many as fit: about 640 of them on an 8 GB card, all of them from 16 GB on by the installer's estimate. While it writes, the engine brings in the experts the answer keeps choosing, and on the 8 GB card the ones it holds serve about 55% of what an answer uses. Your RAM holds the rest, and your processor runs them at the same time. Your SSD holds the model's file and its draft model, which are used where they are, with no second copy.

Google's small draft model guesses the next words; the model checks them all at once and keeps the right ones

  • Guess, then check: Google's small draft model for it, a file of 0.5 GB that the installer downloads with the model, proposes the next few words. The model checks them all at once. You get the same answer, 1.3-1.6x sooner.
  • Long texts are read in big pieces (1,024 tokens at a time). When the card does not hold every expert, the others pass through it once for the whole prompt.

Every part and its numbers: docs/BACKENDS.md.

Credits and license

Qwen3.8-Flash-Next is by the Qwen team. It was compressed by ISTA-DASLab, UkisAI (Swift 1.5) and Unsloth. GLM 5.3 Flash runs from the files of neuralll, Unsloth and patrickbdevaney. Qwen3.6-35B-A3B is by the Qwen team and Gemma 4 by Google, both in Unsloth's quantization. Ornith 1.5 is by ornith-ai, in mudler's APEX quantization. PolyStrata is built on Strata by Niko1221 and its contributors, and uses parts of llama.cpp / ggml. All credits: docs/HOW_IT_WORKS.md. PolyStrata is open source under the MIT License. A few parts and every model have their own licenses (which ones).

About

Run large AI models on your own PC: Qwen3.8-Flash-Next, GLM 5.3 Flash, Qwen3.6, Ornith 1.5 and Gemma 4. One-click install for Windows and Linux, OpenAI/Anthropic API on localhost, image input. Built on Strata.

Topics

Resources

Security policy

Stars

66 stars

Watchers

3 watching

Forks

Releases

Packages

Contributors

Languages