🛠️[v0.4.2] Release Note: RPC Device Selection and Extended Batch / Runtime Bindings #185
JamePeng
announced in
Announcements
Replies: 1 comment
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment


Uh oh!
There was an error while loading. Please reload this page.
[0.4.2] — RPC Device Selection and Extended Batch / Runtime Bindings
This release introduces a major expansion of RPC backend support, allowing
Llamainstances to explicitly select remote and local devices for model execution. RPC endpoints are now validated before registration, process-wide registrations are reused safely, and server-side configuration with prebuilt ggml-rpc-server from the official llama.cpp releases can place model layers entirely on selected remote devices when desired.The release also expands the low-level runtime surface with the extended batch API, multi-tensor GGML buffer allocation bindings, and several ABI corrections across state, MTMD, quantization, optimizer, and callback interfaces. RPC is now enabled across CUDA and Metal wheel builds and verified in Linux, Windows, and macOS CI.
Highlights
RPC device selection
Llamainstance.rpc_local_devices=[].Server RPC support
rpc_local_devicesthrough server model settings.Extended llama.cpp bindings
llama_batch_ext,llama_embd,llama_process_type, andllama_process.GGML backend expansion
Build and CI
ABI fixes
Documentation
API sync
Thanks to @alcoftTAO for RPC backend testing feedback in issue #181.
More information:
6332d8d...6126add
--JamePeng
All reactions