ollama

mirror of https://github.com/ollama/ollama.git synced 2025-09-13 21:02:38 +02:00

Files

Jesse Gross 3fe74fba42 llm: Use first layer as memory buffer in estimation

This is a partial revert of 0478d44 "Fixed over vram allcation dure to
small initial layer sizes."

Previously we used the size of the first layer as an extra reserved
amount of space to buffer our memory estimates. The above commit
changed this to use the largest layer. However, this had performance
impacts on more models than the original commit was trying to fix.

There is just a heuristic without an ideal solution so this goes back
to the historic behavior.

Fixes: #10765, #10756, #10752, #10726

2025-05-19 14:03:34 -07:00

llm_darwin.go

…

llm_linux.go

…

llm_windows.go

win: lint fix (#10571 )

2025-05-05 11:08:12 -07:00

memory_test.go

Move quantization to new backend (#10363 )

2025-05-06 11:20:48 -07:00

memory.go