llama.cpp

mirror of https://github.com/ggml-org/llama.cpp.git synced 2025-10-27 08:21:30 +00:00

Files

65a 4afb0a746f server : Support multimodal completion and embeddings prompts in JSON format (#15108 )

- Use server_tokens in more places in server and util.cpp
- Convert most functions that used llama_tokens to server_tokens
- Modify input tokenizer to handle JSON objects as subprompts
- Break out MTMD prompt parsing into utility function
- Support JSON objects with multimodal_data arrays for MTMD prompts along with other existing types
- Add capability to model endpoint to indicate if client can send multimodal data
- Add tests.

2025-08-22 10:10:14 +02:00

batched-bench

batched-bench : use rand tokens (#15398 )

2025-08-19 08:45:12 +03:00

cvector-generator

llama : deprecate llama_kv_self_ API (#14030 )

2025-06-06 14:11:15 +03:00

export-lora

mtmd : fix 32-bit narrowing issue in export-lora and mtmd clip (#14503 )

2025-07-25 13:08:04 +02:00

gguf-split

scripts : make the shell scripts cross-platform (#14341 )