ollama

mirror of https://github.com/ollama/ollama.git synced 2025-11-11 04:57:52 +01:00

Files

Jesse Gross 1fc35f1260 kvcache: Clean up sliding window state with independent batches

Sliding windows models (e.g. gpt-oss, gemma3) remove tokens that
are out of the cache's window each time we start a new forward pass.

The cache storage needs to handle the window size for each sequence
plus the batch size, since the batch needs to attend to the full
window size. This means that we have greater than a window size
stored while processing the batch.

When the next batch comes, we are currently only looking at the
sequences in the incoming batch to slide the window forward.
However, we also need to clean up the other sequences that might
be occupying space in the batch processing buffer to ensure each
sequence is only using its window size of storage. Failure to do
this can result in "no kv cache slot found" errors.

Fixes: #10127

2025-10-08 16:43:14 -07:00

cache.go

ollamarunner: Preallocate worst case graph at startup

2025-04-08 10:01:28 -07:00

causal_test.go

kvcache: Clean up sliding window state with independent batches

2025-10-08 16:43:14 -07:00

causal.go

kvcache: Clean up sliding window state with independent batches

2025-10-08 16:43:14 -07:00

encoder.go

ollamarunner: Preallocate worst case graph at startup

2025-04-08 10:01:28 -07:00

wrapper.go

ollamarunner: Preallocate worst case graph at startup

2025-04-08 10:01:28 -07:00