一款口袋级多模态大语言模型,在手机上实现超高效图像与视频理解
GitHub | MiniCPM Wiki(中文) | CookBook | Demo | 飞书(Lark)
MiniCPM-V 4.6 Thinking 是 MiniCPM-V 4.6 的长思维链推理版本。它在生成最终答案之前会先输出一条显式的推理轨迹,从而显著提升在复杂多模态推理、数学和 OCR 密集型任务上的表现,同时保持了与 MiniCPM-V 4.6 相同的边缘友好架构(SigLIP2-400M 视觉编码器 + Qwen3.5-0.8B 语言模型)以及混合 4x/16x 视觉 token 压缩策略。
整体性能(Thinking 版本)
高并发吞吐量
单请求 TTFT(毫秒)
MiniCPM-V 4.6 可部署于三大主流端侧平台——iOS、Android 和 HarmonyOS。以下片段为手机设备的原始屏幕录制,未经任何编辑。
| iPhone iPhone 17 Pro Max | Android Redmi K70 | HarmonyOS HUAWEI nova 14 |
![]() | ![]() | ![]() |
pip install "transformers[torch]>=5.7.0" torchvision torchcodec关于 CUDA 兼容性的说明:
torchcodec(用于视频解码)可能与某些 CUDA 版本存在兼容性问题。例如,torch>=2.11默认捆绑 CUDA 13.1,而 CUDA 12.x 环境可能会遇到诸如RuntimeError: Could not load libtorchcodec之类的错误。有两种解决方法:
- 将
torchcodec替换为PyAV— 支持图像和视频推理,且不受 CUDA 版本限制:pip install "transformers[torch]>=5.7.0" torchvision av- 固定 CUDA 版本 在安装 torch 时,使其与你的环境匹配(例如 CUDA 12.8):
pip install "transformers>=5.7.0" torchvision torchcodec --index-url https://download.pytorch.org/whl/cu128
from transformers import AutoModelForImageTextToText, AutoProcessor
model_id = "openbmb/MiniCPM-V-4.6-Thinking"
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForImageTextToText.from_pretrained(
model_id, torch_dtype="auto", device_map="auto"
)
# Flash Attention 2 is recommended for better acceleration and memory saving,
# especially in multi-image and video scenarios.
# model = AutoModelForImageTextToText.from_pretrained(
# model_id,
# torch_dtype=torch.bfloat16,
# attn_implementation="flash_attention_2",
# device_map="auto",
# )messages = [
{
"role": "user",
"content": [
{"type": "image", "url": "https://huggingface.co/datasets/openbmb/DemoCase/resolve/main/refract.png"},
{"type": "text", "text": "What causes this phenomenon?"},
],
}
]
downsample_mode = "16x" # Using `downsample_mode="4x"` for Finer Detail
inputs = processor.apply_chat_template(
messages, tokenize=True, add_generation_prompt=True,
return_dict=True, return_tensors="pt",
downsample_mode=downsample_mode,
max_slice_nums=36,
).to(model.device)
generated_ids = model.generate(**inputs, downsample_mode=downsample_mode, max_new_tokens=512)
generated_ids_trimmed = [
out_ids[len(in_ids):] for in_ids, out_ids in zip(inputs.input_ids, generated_ids)
]
output_text = processor.batch_decode(
generated_ids_trimmed, skip_special_tokens=True, clean_up_tokenization_spaces=False
)
print(output_text[0])messages = [
{
"role": "user",
"content": [
{"type": "video", "url": "https://huggingface.co/datasets/openbmb/DemoCase/resolve/main/football.mp4"},
{"type": "text", "text": "Describe this video in detail. Follow the timeline and focus on on-screen text, interface changes, main actions, and scene changes."},
],
}
]
downsample_mode = "16x" # Using `downsample_mode="4x"` for Finer Detail
inputs = processor.apply_chat_template(
messages, tokenize=True, add_generation_prompt=True,
return_dict=True, return_tensors="pt",
downsample_mode=downsample_mode,
max_num_frames=128,
stack_frames=1,
max_slice_nums=1,
use_image_id=False,
).to(model.device)
generated_ids = model.generate(**inputs, downsample_mode=downsample_mode, max_new_tokens=2048)
generated_ids_trimmed = [
out_ids[len(in_ids):] for in_ids, out_ids in zip(inputs.input_ids, generated_ids)
]
output_text = processor.batch_decode(
generated_ids_trimmed, skip_special_tokens=True, clean_up_tokenization_spaces=False
)
print(output_text[0])您可以通过向 apply_chat_template 传递额外参数来自定义图像/视频处理:
| 参数 | 默认值 | 适用范围 | 说明 |
|---|---|---|---|
downsample_mode | "16x" | 图像与视频 | 视觉令牌降采样。"16x" 合并令牌以提高效率;"4x" 保留 4 倍令牌以呈现更精细的细节。该参数同样需要传递给 generate()。 |
max_slice_nums | 9 | 图像与视频 | 高分辨率图像分割时的最大切片数量。数值越大,对大图保留的细节越多。建议:图像设为 36,视频设为 1。 |
max_num_frames | 128 | 仅视频 | max_num_frames 参数动态控制时间上下文长度,并防止显存溢出: 短视频(时长 ≤ max_num_frames 秒):处理器默认采用 1 FPS 采样,逐秒捕捉细节而不会触及上限。 长视频(时长 > max_num_frames 秒):处理器自动切换为均匀采样,在整个时间轴上均匀选取恰好 max_num_frames 帧。 |
stack_frames | 1 | 仅视频 | 每秒总采样点数。1 = 仅主帧(无堆叠)。N(N>1)= 每秒 1 个主帧 + N−1 个子帧;子帧合成为网格图像并与主帧交错排列。建议短视频设为 1,长视频设为 3 或 5。 |
use_image_id | True | 图像与视频 | 是否在每个图像/帧占位符前添加 <image_id>N</image_id> 标签。图像设为 True,视频设为 False。 |
注意:
downsample_mode必须同时传递给apply_chat_template(以确保占位符数量正确)和generate(供视觉编码器使用)。其余参数仅需传递给apply_chat_template。
transformers serve 部署服务 Hugging Face Transformers 内置了轻量级的 OpenAI 兼容服务器,适用于快速测试和中负载部署。
pip install "transformers[serving]>=5.7.0"启动服务器:
transformers serve openbmb/MiniCPM-V-4.6-Thinking --port 8000 --host 0.0.0.0 --continuous-batching发送请求:
curl -s http://localhost:8000/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "openbmb/MiniCPM-V-4.6-Thinking",
"messages": [{
"role": "user",
"content": [
{"type": "image_url", "image_url": {"url": "https://huggingface.co/datasets/openbmb/DemoCase/resolve/main/refract.png"}},
{"type": "text", "text": "What causes this phenomenon?"}
]
}]
}'工具调用示例:
curl -s http://localhost:8000/v1/chat/completions -H 'Content-Type: application/json' -d '{
"model": "openbmb/MiniCPM-V-4.6-Thinking",
"messages": [{"role": "user", "content": [
{"type": "text", "text": "the weather of Beijing"}
]}],
"tools": [{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get the current weather for a given location",
"parameters": {
"type": "object",
"properties": {
"location": {"type": "string", "description": "City name"}
},
"required": ["location"]
}
}
}]
}'该模型会先返回一段自然语言解释,随后在内容字段中嵌入结构化的<tool_call>代码块。需要注意的是,目前transformers库中尚未添加针对此格式的专用工具调用解析器,因此现阶段需要通过正则表达式手动提取工具调用。
{
"id": "f4f09c7d-8045-4cb1-ade9-07aa5dee637d",
"choices": [
{
"finish_reason": "stop",
"index": 0,
"message": {
"content": "I need to check the current weather for Beijing, so I will call the get_weather function.\n\n<tool_call>\n<function=get_weather>\n<parameter=location>\nBeijing\n</parameter>\n</function>\n</tool_call>",
"role": "assistant"
}
}
],
"created": 1778748859,
"model": "openbmb/MiniCPM-V-4.6-Thinking@main",
"object": "chat.completion",
"usage": {
"completion_tokens": 47,
"prompt_tokens": 283,
"total_tokens": 330
}
}在某些情况下,模型可能会将转义换行符 \n 以字符串字面量的形式输出,而非实际的换行。为了在界面层正确渲染文本,您可以使用以下工具函数。该函数会谨慎地将字面量 \n 替换为真正的换行,同时保护那些 \n 具有特定语义的场景。
工具函数:
import re
_PATTERN = re.compile(
r'(```[\s\S]*?```' # fenced code blocks
r'|`[^`]+`' # inline code
r'|\$\$[\s\S]*?\$\$' # display math
r'|\$[^$]+\$' # inline math
r'|\\$[\s\S]*?\\$' # $...$
r'|\\
$$[\s\S]*?\\$$
' #
$$...$$
r')'
r'|(?<!\\)(?:\\r\\n|\\[nr])'
)
def normalize_response_text(text: str) -> str:
"""
Lightweight post-processing: Converts literal '\\n' to actual newlines,
while protecting code blocks, inline code, and LaTeX commands.
"""
if not isinstance(text, str) or "\\" not in text:
return text
return _PATTERN.sub(lambda m: m.group(1) or '\n', text)我们已将 MiniCPM-V 4.6 适配至 iOS、Android 和 HarmonyOS 平台,所有边缘适配代码均已完全开源。开发者只需几步即可复现端侧部署体验。请访问我们的边缘部署仓库查看各平台的构建指南,或前往下载页面直接体验预构建应用。
MiniCPM-V 4.6 支持多种推理与训练框架。以下是各框架的快速上手命令。完整细节请参阅我们的示例手册。
vllm serve openbmb/MiniCPM-V-4.6-Thinking \
--port 8000 \
--enable-auto-tool-choice \
--tool-call-parser qwen3_coder \
--default-chat-template-kwargs '{"enable_thinking": true}'注意:
--enable-auto-tool-choice与--tool-call-parser qwen3_coder用于启用工具/函数调用功能。如果无需使用工具,可省略这些参数,直接运行vllm serve openbmb/MiniCPM-V-4.6-Thinking即可。
curl -s http://localhost:8000/v1/chat/completions -H 'Content-Type: application/json' -d '{
"model": "openbmb/MiniCPM-V-4.6-Thinking",
"messages": [{"role": "user", "content": [
{"type": "image_url", "image_url": {"url": "https://huggingface.co/datasets/openbmb/DemoCase/resolve/main/refract.png"}},
{"type": "text", "text": "What causes this phenomenon?"}
]}]
}'工具调用示例:
curl -s http://localhost:8000/v1/chat/completions -H 'Content-Type: application/json' -d '{
"model": "openbmb/MiniCPM-V-4.6-Thinking",
"messages": [{"role": "user", "content": [
{"type": "text", "text": "北京的天气"}
]}],
"tools": [{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get the current weather for a given location",
"parameters": {
"type": "object",
"properties": {
"location": {"type": "string", "description": "City name"}
},
"required": ["location"]
}
}
}]
}'python -m sglang.launch_server --model openbmb/MiniCPM-V-4.6-Thinking --port 30000curl -s http://localhost:30000/v1/chat/completions -H 'Content-Type: application/json' -d '{
"model": "openbmb/MiniCPM-V-4.6-Thinking",
"messages": [{"role": "user", "content": [
{"type": "image_url", "image_url": {"url": "https://huggingface.co/datasets/openbmb/DemoCase/resolve/main/refract.png"}},
{"type": "text", "text": "What causes this phenomenon?"}
]}]
}'llama-server -m MiniCPM-V-4.6-Q4_K_M.gguf --port 8080curl -s http://localhost:8080/v1/chat/completions -H 'Content-Type: application/json' -d '{
"model": "MiniCPM-V-4.6",
"messages": [{"role": "user", "content": [
{"type": "image_url", "image_url": {"url": "https://huggingface.co/datasets/openbmb/DemoCase/resolve/main/refract.png"}},
{"type": "text", "text": "What causes this phenomenon?"}
]}]
}'llamafactory-cli train examples/train_lora/minicpmv4_6_lora_sft.yamlswift sft --model_type minicpm-v-4_6 --dataset <your-dataset>👏 欢迎探索 MiniCPM-o/V 的关键技术以及我们团队的其他多模态项目:
技术报告: MiniCPM-o 4.5 | MiniCPM-V 4.5 | MiniCPM-o 2.6 | MiniCPM-Llama3-V 2.5 | MiniCPM-V 2.0
其他多模态项目: VisCPM | RLPR | RLHF-V | LLaVA-UHD | RLAIF-V | LLaVA-UHD-v4
如果您觉得我们的模型、代码或论文对您有帮助,欢迎引用我们的论文 📝 并为我们点亮星标 ⭐️!
@proceedings{yu2025minicpmv45cookingefficient,
title={MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe},
author={Tianyu Yu and Zefan Wang and Chongyi Wang and Fuwei Huang and Wenshuo Ma and Zhihui He and Tianchi Cai and Weize Chen and Yuxiang Huang and Yuanqian Zhao and others},
year={2025},
url={https://arxiv.org/abs/2509.18154},
}
@article{yao2024minicpm,
title={MiniCPM-V: A GPT-4V Level MLLM on Your Phone},
author={Yao, Yuan and Yu, Tianyu and Zhang, Ao and Wang, Chongyi and Cui, Junbo and Zhu, Hongji and Cai, Tianchi and Li, Haoyu and Zhao, Weilin and He, Zhihui and others},
journal={arXiv preprint arXiv:2408.01800},
year={2024}
}