Skip to content

Commit 169d5e1

Browse files
committed
Bump version to 0.3.41
- This release is mainly focused on improving the multimodal architecture, making MTMD chat handling cleaner, more flexible, and easier to extend for future image, audio, and video workflows. I also continued syncing with the latest llama.cpp APIs and improved several developer-facing diagnostics around model templates, shared library loading, and Windows OpenMP runtime discovery. Signed-off-by: JamePeng <jame_peng@sina.com>
1 parent b9b5859 commit 169d5e1

2 files changed

Lines changed: 134 additions & 1 deletion

File tree

‎CHANGELOG.md‎

Lines changed: 133 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -7,6 +7,139 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
77

88
## [Unreleased]
99

10+
## [0.3.41] Template-Driven MTMD, Broader Multimodal Inputs, and Smarter N-Gram Drafting
11+
12+
- refactor(mtmd): extract prompt rendering and media marker normalization
13+
- Add extra_template_arguments to MTMD chat handlers and pass them through to the Jinja chat template render call. This allows generic model templates to receive render-time options such as enable_thinking, add_vision_id, or model-specific template jinja variables.
14+
- Extract MTMD prompt rendering into dedicated helpers:
15+
* _render_mtmd_prompt() for pure chat template rendering
16+
* _replace_media_placeholders() for normalizing rendered media tags and URLs into the MTMD runtime marker
17+
* _render_and_replace_media() for the combined render-and-normalize stage
18+
- This removes inline render/replace logic from _process_mtmd_prompt(), keeps media marker validation after normalization, and improves separation between prompt construction and MTMD tokenization.
19+
20+
- docs(README): add GenericMTMDChatHandler usage guide
21+
- Replace the legacy Llava multimodal loading example with a GenericMTMDChatHandler
22+
usage guide for template-driven multimodal GGUF models.
23+
- Document loading mmproj through Llama, chat template resolution order,
24+
extra_template_arguments for model-specific Jinja variables, and when to prefer
25+
a dedicated multimodal chat handler.
26+
- Also clarify the mmproj_path naming, llama_multimodal migration, and note that
27+
the generic handler is intended as a flexible fallback for models without
28+
dedicated handlers and may require additional testing for model-specific
29+
prompting behavior.
30+
- Update Generic MTMD Chat Handler directory index.
31+
32+
- refactor(mtmd): extract mtmd_tokenize into _mtmd_tokenize standalone helper
33+
- Introduce `_mtmd_tokenize()` to encapsulate llama.cpp mtmd_tokenize binding
34+
- Decouple hybrid tokenization logic from `_process_mtmd_prompt`
35+
- Improve separation of concerns between prompt construction and C++ binding
36+
- Preserve strict media marker validation to ensure token/bitmap alignment
37+
38+
- feat(speculative): Improve ngram-map draft selection and accept feedback
39+
- Store accepted draft lengths per key/value and truncate future drafts accordingly
40+
- Make key-only mode draft on any key match without applying min_hits
41+
- Select k4v continuations by frequency instead of latest occurrence
42+
- Skip ambiguous k4v drafts when the top continuation is not dominant
43+
- Track fixed-size k4v continuations to keep frequency statistics comparable
44+
45+
- feat(mtmd): broaden multimodal media extraction
46+
- Broaden MTMD media extraction to support common multimodal content shapes used
47+
by model chat templates.
48+
- In addition to OpenAI-style image_url/audio_url/video_url chunks, accept
49+
image/audio/video typed chunks and direct media keys such as {"image": "..."},
50+
{"audio": "..."}, or {"video": "..."}. This keeps the extracted media list
51+
aligned with templates that emit placeholders for image, audio, or video content
52+
without requiring URL-specific chunk names.
53+
- Add a shared helper for extracting URLs, local paths, existing data URIs, or
54+
inline base64 payloads from media content items. Preserve capability checks,
55+
strict input_audio format validation, and explicit errors for missing or
56+
ambiguous media payloads.
57+
58+
- feat(mtmd): enhance generic chat template support
59+
- Enhance GenericMTMDChatHandler to better support model-provided chat templates.
60+
- Allow the generic handler to accept an optional named chat template, load it
61+
from the model at call time via llama_model_chat_template(), fall back to the
62+
model's default chat template, and finally use the built-in MTMD CHAT_FORMAT
63+
when no model template is available.
64+
- Also expand the generic media placeholder list for common multimodal templates
65+
and document the handler as a template-driven MTMD implementation. This prepares
66+
the generic path for a later render-driven placeholder replacement pass.
67+
68+
- fix(model): handle missing chat templates
69+
- Update `LlamaModel.model_chat_template()` to return Optional[str] and accept
70+
name=None for the default model chat template.
71+
- `llama_model_chat_template()` may return nullptr when no chat template is
72+
available. Handle that case explicitly instead of decoding a null pointer, and
73+
return None so callers can apply their own fallback logic.
74+
75+
- fix(vocab): update `LlamaModel.vocab_type` to use self.vocab and add None checks
76+
77+
- refactor(mtmd): move multimodal handlers to separate module `llama_multimodal`
78+
- Move `MTMDChatHandler`, `GenericMTMDChatHandler``, and model-specific multimodal
79+
chat handlers out of `llama_chat_format.py` into `llama_multimodal.py`.
80+
- `llama_chat_format.py` has grown too large and difficult to maintain, especially
81+
as MTMD support expands beyond image-only use cases. Splitting multimodal
82+
handling into its own module makes the chat formatting layer smaller and keeps
83+
media loading, MTMD tokenization, multimodal KV-cache bookkeeping, and handler
84+
implementations in a dedicated place.
85+
- This also prepares the codebase for broader multimodal support and future video
86+
frame / image batch evaluation, where the media-processing path will need to
87+
evolve independently from text-only chat formatting.
88+
- Keep backward-compatible re-exports from `llama_chat_format.py` so existing
89+
imports continue to work.
90+
- Also keep `clip_model_path` as a deprecated initialization alias for
91+
`mmproj_path` in the base MTMD handler.
92+
- docs: update mtmd chat handler import paths in README
93+
- Update import statements for multi-modal chat handlers from llama_cpp.llama_chat_format to llama_cpp.llama_multimodal in the documentation examples.
94+
95+
- feat: Implemented generic multimodal chat handler prototype (by **@alcoftTAO**)
96+
97+
- docs(README): Added command prompt scenario for README.md (by **@patrikpatrik**)
98+
- Updated command prompt scenario under Configuration -> Environment Variables
99+
- Sanity checking after successful installation of wheel
100+
101+
- feat(MTMDChatHandler): add chunk type helpers
102+
- Add small helper methods `_is_text_chunk`/`_is_image_chunk`/`_is_audio_chunk` for checking
103+
MTMD text, image, and audio chunk type enum values.
104+
- This keeps MTMD prompt processing easier to read and avoids repeating direct
105+
enum comparisons when building token spans for text and media chunks.
106+
107+
- feat(mtmd): add video input support to `MTMDChatHandler`
108+
- Add video_url handling to the MTMD chat template and media extraction
109+
pipeline. Detect whether the loaded libmtmd build supports video helpers
110+
and reject video inputs early when MTMD_VIDEO is unavailable.
111+
- Update media loading and bitmap creation for the new helper wrapper API.
112+
mtmd_helper_bitmap_init_from_buf now returns a bitmap wrapper containing
113+
both the decoded bitmap and an optional video helper context, so keep the
114+
video context alive until mtmd_tokenize completes and release it afterward.
115+
- Also consolidate duplicated audio/video byte loading into a shared
116+
_load_bytes helper, reuse it for image loading, and add richer default HTTP
117+
headers for remote media requests.
118+
119+
- build(CMakelists): Improve Windows LLVM OpenMP runtime `libomp140.x86_64.dll` discovery
120+
- Also improve diagnostics by reporting the selected runtime source and path,
121+
warning when an explicit override points to a missing file, and keeping a clear
122+
runtime warning when no OpenMP DLL can be found.
123+
- prefer VS 2022 VC143 OpenMP redist and keep System32 as final fallback。
124+
125+
- feat(_ctypes_extensions): improve error diagnostics for shared library loading
126+
When `load_shared_library` fails, the resulting `RuntimeError` now
127+
includes a listing of the contents of the searched directories. This
128+
provides immediate context to help developers diagnose missing, misplaced,
129+
or incorrectly named library files.
130+
131+
- Added `_format_library_dir_contents` to safely format directory listings.
132+
- Appended the directory listing to the failure message.
133+
- Confined this diagnostic work strictly to the failure path to avoid any
134+
performance overhead during successful imports.
135+
136+
- feat: Update llama.cpp to [ggml-org/llama.cpp/commit/3899b39ce2acc2e019f149b7107f24b6ca297390](https://github.com/ggml-org/llama.cpp/commit/3899b39ce2acc2e019f149b7107f24b6ca297390)
137+
138+
- feat: Sync llama.cpp llama/mtmd/ggml API Binding 20260707
139+
140+
More information see: https://github.com/JamePeng/llama-cpp-python/compare/12861b918f67b62f78f28c5cabb7223f766e1097...b9b58594023ab673c2dda6723f8909d85d65a2e5
141+
142+
10143
## [0.3.40-Milestone] Reasoning Budget Control, Gemma 4 12B Support, Enhanced Jinja2ChatFormatter, NGram k/k4v Speculative Decoding, Faster Native Sampling and Multimodal Improvements
11144

12145
- feat(internals): Add `ReasoningBudgetSampler` support

‎llama_cpp/__init__.py‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -1,4 +1,4 @@
11
from .llama_cpp import *
22
from .llama import *
33

4-
__version__ = "0.3.40"
4+
__version__ = "0.3.41"

0 commit comments

Comments
 (0)