ollama

mirror of https://github.com/ollama/ollama.git synced 2025-04-18 00:21:20 +02:00

Author	SHA1	Message	Date
Michael Yang	7632c4cb3c	fix tests	2025-04-03 09:06:39 -07:00
Michael Yang	7946618dc3	metal: op_neg	2025-04-02 16:55:06 -07:00
Michael Yang	2ab14468a8	ml: add repeat op repeat is a convenience operation for repeating a tensor n times along a dimention. this can replace instances where the same tensors are stacked together	2025-04-01 15:25:20 -07:00
Michael Yang	6184028fc0	fix convert	2025-03-31 13:12:34 -07:00
Michael Yang	557c641697	2d rope	2025-03-31 10:34:17 -07:00
Michael Yang	863ba57477	fixes	2025-03-25 13:57:24 -07:00
jmorganca	4586e137fe	wip	2025-03-23 21:41:18 -07:00
jmorganca	8dd2a81f8c	wip	2025-03-22 22:33:39 -07:00
Michael Yang	74bd09652d	ml/backend/ggml: load tensors in 32KiB chunks	2025-03-21 14:43:52 -07:00
Jesse Gross	0ff28758b3	ollamarunner: Provide mechanism for backends to report loading progress This enables the runner to report progress back to the Ollama server, both for showing status to the user and also to prevent the server from killing the runner if it thinks things have stalled. Most of the infrastructure was already there, this extends it to be available to the backends.	2025-03-21 10:44:26 -07:00
Bruce MacDonald	df94175a0f	ggml: return error on failure to read tensor data (#9872 ) When converting a ggml model if there is a failure to read tensor data a nil error value was being returned. It should be assigned to the actual error from reading.	2025-03-18 16:51:33 -07:00
Michael Yang	021dcf089d	Merge pull request #9824 from ollama/mxyng/sched conditionally enable parallel pipelines	2025-03-17 15:41:37 -07:00
Jeffrey Morgan	364629b8d6	ml/backend/ggml: allocate memory with malloc when loading model (#9822 )	2025-03-17 13:32:40 -07:00
Michael Yang	4561fff36e	conditionally enable parallel pipelines	2025-03-17 09:46:07 -07:00
shane.xb.qian	30d7a59ba8	ollama-debug.c: change 'ld' to 'PRIi64' * macOS has different definition per info from @mxyng	2025-03-13 17:10:37 +08:00
shane.xb.qian	85ab552028	ollama-debug.c: correct mistype Signed-off-by: shane.xb.qian <shane.qian@foxmail.com>	2025-03-12 22:32:30 +08:00
Michael Yang	63a394068c	use 2d pooling	2025-03-11 14:49:20 -07:00
Michael Yang	c5cbe4fc2a	fallback to cpu	2025-03-11 14:49:19 -07:00
Michael Yang	9e4642e9b3	ollama debug tensor	2025-03-11 14:49:19 -07:00
Michael Yang	6b0486c216	duplicate token_embd to output	2025-03-11 14:49:19 -07:00
Michael Yang	8934324b72	use fast attention	2025-03-11 14:49:18 -07:00
Michael Yang	0df1800436	set non-causal attention	2025-03-11 14:49:18 -07:00
Michael Yang	4b037a97dc	add gemma vision encoder	2025-03-11 14:49:17 -07:00
Patrick Devine	5f74d1fd47	gemma2 impl	2025-03-11 14:35:08 -07:00
Michael Yang	9926eae015	fix: pad tensor item if ge zero this produces a nicer output since both positive and negative values produces the same width	2025-03-10 16:18:12 -07:00
Jesse Gross	4100ed7bdd	ml: Add support for quantized KV cache Similar to the llama engine, quantizing the KV cache requires flash attention to be enabled through the Ollama server.	2025-03-07 18:43:39 -08:00
Jesse Gross	25f9b152f9	ggml-backend: Ensure allocation meet backend requirements Backends can impose additional alignment requirements on buffer sizes. We should ensure that we meet these or allocations can fail.	2025-03-07 18:43:39 -08:00
Jesse Gross	98272fbd58	additional review comments	2025-03-07 14:08:21 -08:00
Michael Yang	b27e8f3f10	ml/backend/ggml: use backend buffer type this ensures the tensor is created on the right buffer type for backends such as cpu	2025-03-07 14:08:21 -08:00
Michael Yang	45df786f09	comments	2025-03-07 14:08:21 -08:00
Michael Yang	daaf42e4a4	ml/backend/ggml: clean up	2025-03-07 14:08:21 -08:00
Michael Yang	2dc60d4620	ml/backend/ggml: offload vision to cpu temporary until tensor loading can accurately account for vision models	2025-03-07 14:08:21 -08:00
Michael Yang	b5312f30e8	ml/backend/ggml: handle tensor split	2025-03-07 14:08:21 -08:00
Michael Yang	26c2e0bd35	ml/backend/ggml: handle user specified cpu offloading	2025-03-07 14:08:21 -08:00
Michael Yang	bf920883d5	ml/backend/ggml: set cpu n_threads	2025-03-07 14:08:21 -08:00
Michael Yang	7bae7fa5ce	ml/backend/ggml: create tensor on specific backend some tensors should be created on specific backends to reduce number of copies and improve performance	2025-03-07 14:08:21 -08:00
Michael Yang	764e199d67	kvcache: create cache ctx per layer each cache layer creates and maintains its own context instead of using a large context for all layers	2025-03-07 14:08:21 -08:00
Michael Yang	bfce55db3d	model: load non-repeated tensors into multiple backends some tensors are expected to be used in repeating layers but are not themselves repeated. this change copies these tensors into the same backends as their repeating counterparts to minimize copying tensors between backends	2025-03-07 14:08:21 -08:00
Michael Yang	bab6f34dc0	ml/backend/ggml: update model loading for hybrid/multi backends use a similar strategy as llama.cpp for deciding where tensors should be allocated. this will be improved later to be aware of usable memory before assigning the tensor	2025-03-07 14:08:21 -08:00
Jeffrey Morgan	4289c74359	llama: fix kv loading on snowflake-arctic-embed models (#9536 )	2025-03-07 09:25:34 -08:00
Michael Yang	05a01fdecb	ml/backend/ggml: consolidate system info logging - output backend system info when initializing the backend. this ensures this information is always present without needing to be called explicitly - convert to structured logging - enumerate devices rather than backends since devices are ordered - track device indices grouped by device name	2025-03-04 15:14:31 -08:00
Michael Yang	ba7d31240e	fix: own lib/ollama directory expand backend loading error handling to catch more problems and log them instead of panicing	2025-03-03 13:01:18 -08:00
Jesse Gross	21aa666a1e	ml: Enable support for flash attention The GGML flash attention kernel has specific requirements for padding and permutation. This adds support to the KV cache for conforming to these requirements so that flash attention can be enabled. Flash attention can be used in the same situations as the llama engine and is enabled by the user in the same way.	2025-03-01 20:53:23 -08:00
Jesse Gross	ee141cc821	ml: Empty tensor constructor for tensors In cases where we allocate a tensor and then fully overwrite it with copied data, it is wasteful to first zero out the memory.	2025-03-01 20:53:23 -08:00
Jesse Gross	55e5776c44	ggml-backend: Store parent backend as part of tensor It can be important for a tensor to know what backend it came from - for example, to know if flash attention is enabled.	2025-03-01 20:53:23 -08:00
Jesse Gross	854a9195f3	attention: Remove unnecessary contiguous operations Prior to performing attention, we need to permute query, key and value. Currently we call Contiguous after each of these permutations, which is correct but expensive. Avoiding the 3 calls to Contiguous increases performance by over 20%. The permutations of query and key do not violate the continuity rules for mulmat and the Contiguous call can be simply removed. Value requires a different permutation and does require Contiguous. However, we can use the copy into the cache as a way to perform this without further overhead. To support this and avoid unexpected tensor shapes that are seen by models, we need tighter integration between attention, cache and backend. Future optimization will also likely need this structure - for example, flash attention has special padding requirements in the cache and other backends may have their own needs. This further contains the operations that go into attention so that these and other optimizations can be handled transparently. Models that have special requirements for attention can still implement their own version of it.	2025-03-01 20:53:23 -08:00
Michael Yang	3e8b8a1933	ml: update Context.Forward interface update Context.Forward to accept multiple tensors to match Context.Compute signature update Context.Forward to return Context such that it can be chained with Context.Compute	2025-02-27 22:27:16 +00:00
Michael Yang	53d2990d9b	model: add bos token if configured	2025-02-27 21:04:59 +00:00
Michael Yang	a59f665235	ml/backend/ggml: fix debug logging	2025-02-27 18:30:57 +00:00
Jeffrey Morgan	a5272130c4	ml/backend/ggml: follow on fixes after updating vendored code (#9388 ) Fixes sync filters and lowers CUDA version to 11.3 in test.yaml	2025-02-26 22:33:53 -08:00

1 2

79 Commits