VLM Benchmark

modelBPU
core num
qtypemax contextTTFT
(ms)
Decode
(TPS)
memory
(GB)
Qwen2.5-VL-
3B-Instruct
(448*448)
vit 4
prefill 4
decode 4
vit w8
prefill w4
decode w4
1024vit 39.1
prefill 86.2
all 125.3
70.43.5
Qwen2.5-VL-
3B-Instruct
(448*448)
vit 4
prefill 4
decode 4
vit w8
prefill w8
decode w8
1024vit 38.5
prefill 94.8
all 133.3
47.15.1
Qwen2.5-VL-
7B-Instruct
(448*448)
vit 4
prefill 4
decode 4
vit w8
prefill w4
decode w4
1024vit 39.1
prefill 136.6
all 175.7
40.56.1
Qwen3-VL-
2B-Instruct
(448*448)
vit 4
prefill 4
decode 4
vit w8
prefill w4
decode w4
1024vit 27.0
prefill 66.3
all 93.3
95.93.6
Qwen3-VL-
2B-Instruct
(448*448)
vit 4
prefill 4
decode 4
vit w8
prefill w8
decode w8
1024vit 26.6
prefill 69.2
all 95.8
71.34.2
Qwen3-VL-
4B-Instruct
(448*448)
vit 4
prefill 4
decode 4
vit w8
prefill w4
decode w4
1024vit 26.7
prefill 139.3
all 166.0
53.45.8
Qwen3-VL-
4B-Instruct
(448*448)
vit 4
prefill 4
decode 4
vit w8
prefill w8
decode w8
1024vit 26.7
prefill 141.6
all 168.3
35.87.4
Qwen3-VL-
8B-Instruct
(448*448)
vit 4
prefill 4
decode 4
vit w8
prefill w4
decode w4
1024vit 37.2
prefill 198.5
all 235.7
35.18.1
InternVL2-2B
(448*448)
vit 4
prefill 4
decode 4
vit w8
prefill w4
decode w4
1024vit 32.0
prefill 30.0
all 62.0
119.12.0

Annotation:

  • TTFT(ms): The Time to First Token (TTFT) of the VLM is the sum of the ViT processing time and the prefill processing time.