|
Literature
Africa |
LLMOllamaThe Ollama model runner is launched inside a docker container on the main workstations (Didingwe, Odin, Ukhozi). We use the Open-WebUI image to provide and version the web front-end and ollama server API together. Models are stored on the scratch disk of each of the machines. ColibriWe are testing Colibrė on Odin right now. This is due to the model sizes, Odin has a 2TB scratch drive vs 1TB on Didingwe and Ukhozi (though it is SATA not M.2 - we don't have any disk benchmarks for these). Tim has downloaded colibri in home area from the tarball on the release page: https://github.com/JustVugg/colibri/releases
To get a model, we seem to need the Huggingface CLI.
This can be installed with ./coli info stated no model directory given. pass --model <dir>, or set COLI_MODEL=<dir>" Running xXXXx x colibri v1.12.1
xxxxXXXXxXX tiny engine, immense model
XXXXXXX GLM-5.2/5.3 · 744B MoE · 299 GB on disk
XXXX info
X
----------------------------------------------------------
model /mnt/scratch/brooks/huggingface/cache/hub/models--mastouri--GLM-5.2-colibri-E8-IQ3-with-int8-mtp/snapshots/a5b64ba8fc03949c404d7c52d281f50672ba095a
arch hidden 6144 · 78 layer · 256 expert/layer · top-8
shards 141 files · 299 GB on disk
RAM 527 GB total · 455.7 GB available
disk 882 GB free
engine ready
|