ollama

mirror of https://github.com/ollama/ollama.git synced 2025-03-20 14:52:59 +01:00

History

Daniel Hiltgen 34b9db5afc Request and model concurrency

This change adds support for multiple concurrent requests, as well as
loading multiple models by spawning multiple runners. The default
settings are currently set at 1 concurrent request per model and only 1
loaded model at a time, but these can be adjusted by setting
OLLAMA_NUM_PARALLEL and OLLAMA_MAX_LOADED_MODELS.

2024-04-22 19:29:12 -07:00

bytes.go

Request and model concurrency

2024-04-22 19:29:12 -07:00

format.go

instead of static number of parameters for each model family, get the real number from the tensors (#1022 )

2023-11-08 17:55:46 -08:00

time_test.go

go fmt

2023-10-19 09:21:51 -07:00

time.go

cleanup format time

2023-10-11 11:09:27 -07:00