ollama

mirror of https://github.com/ollama/ollama.git synced 2025-04-03 09:29:49 +02:00

History

Jesse Gross e5bcc51ae1 ggml-backend: Don't recreate the scheduler for each context

We don't need to create and destroy the GGML scheduler for every
context. This introduces extra CPU overhead for every forward
pass and extra memory for contexts that don't actually get scheduled
(for example, KV caches). We can instead just have one scheduler
for the backend and reset it each time we call Compute.

This improves token generation performance by 1-2% and removes
scheduler create/destroy from profile traces.

2025-02-20 14:49:47 -08:00

backend

ggml-backend: Don't recreate the scheduler for each context

2025-02-20 14:49:47 -08:00

next ollama runner (#7913 )

2025-02-13 16:31:21 -08:00

backend.go

ollamarunner: Pass runner performance parameters to backends

2025-02-20 13:27:57 -08:00