distributed-llm

Distributed LLM inference — pool GPUs across multiple devices to run models no single machine can handle


Keywords
llm, inference, distributed, pipeline-parallelism, gpu-sharing, peer-to-peer, open-source
License
Apache-2.0
Install
pip install distributed-llm==0.4.1