- [√] supporting vLLM for text inference
- [√] supporting vLLM for text & image inference
The Multimodal EDA Benchmarking project aims to provide a framework for evaluating and comparing the performance of different models on multimodal data. This project supports various data formats and models, helping researchers and developers better understand the capabilities of models in handling multimodal data.
GPT
3.5-Turbo (text-only)
4 (base64)
4-Turbo (base64)
4v (base64)
4o (base64)
Reka
Core (base64)
Flash (base64)
Edge (base64)
Gemini
1.5 Pro (text-only)
Claude3.5
Sonnet (text-only)
MiniGPT4
Vicuna-7B (image-encoder)
QWen
2-0.5B-Instruct (text-only)
2-7B-Instruct (text-only)
2-72B-Instruct (text-only)
2.5-7B-Instruct (text-only)
VL (image-encoder)
VL-Chat (image-encoder)
InternLM
Chat-20B (text-only)
2-Chat-7B (text-only)
2.5-Chat-7B (text-only)
XComposer-VL-7B (image-encoder)
InternVL
2-8B (image-encoder)
2-40B (image-encoder)
InstructBLIP
Flan-T5-XL (image-encoder)
Flan-T5-XXL (image-encoder)
Blip2
Flan-T5-XL (image-encoder)
Flan-T5-XXL (image-encoder)
LlaMa
2-7B-Chat-HF (text-only)
2-13B-Chat-HF (text-only)
3-8B-Instruct (text-only)
3.1-8B-Instruct (text-only)
3.2-11B-Vision-Instruct (image-encoder)
3.2-90B-Vision-Instruct (image-encoder)
MiniCPM
1B-SFT (text-only)
2B-SFT (text-only)
3-4B (text-only)
V (image-encoder)
V2 (image-encoder)
LlaMa-3-V2.5 (image-encoder)
YiVL
6B (image-encoder)
34B (image-encoder)
YiChat
6B (text-only)
ChatGLM3
6B (text-only)
Kosmos
2 (image-encoder)
- reka_flash
- reka_edge
- blip2_flan_t5_xl
- blip2_flan_t5_xxl
- instructblip_flan_t5_xl
- kosmos2_patch14_224
- minicpm2b_sft_bf16
- minicpm3_4b
- minicpm_llama3_v_2_5
- chatglm3_6b
- gpt4
- yi_chat_6b
- yi_vl_34b
- llama3_8b_instruct
- llama3_1_70b_instruct (no general)
- llama3_2_11b_vision_instruct
- intern_chat_20b
- internvl_mini_chat_2b_v1_5 (no general)
- internvl2_40b
- internlm_xcomposer_vl_7b
- qwen_2_7b_instruct
- gpt4o
- gpt4_turbo
- gpt4v
- MiniGPT4-vicuna7b (no general)
- Llama 2-7b
- Llama 2-13b
- Internlm chat 7b
- InternVL-Chat-V1.5
- InternLM2 chat 7b
- InternLM2 Chat 20b
- InternLM2.5 Chat 7b
- InternLM2.5 Chat 20b
- Qwen2 72b instruct
- QWen 2.5 7b instruct
- QWen 2.5 72b instruct
- MiniCPM-V
- GPT-3.5-turbo
- QWen2 0.5b Instruct
- internvl2-8b
- instructBlip flan T5 xxl
- Llama3.1-8b-instruct
- miniCPM-V2
- yi vl 6b
# 2024.11.19
# python main.py --field rtl --model reka_core
# python main.py --field general --model reka_flash (Done)
# python main.py --field general --model reka_edge (Done)
# python main.py --field general --model blip2_flan_t5_xl (Done)
# python main.py --field general --model blip2_flan_t5_xxl (Done)
# python main.py --field general --model gemini_1_5_pro
# python main.py --field general --model claude_3_sonnet
# python main.py --field general --model instructblip_flan_t5_xl (Done)
# python main.py --field general --model kosmos2_patch14_224 (Done)
# python main.py --field general --model minicpm1b_sft_bf16
# python main.py --field general --model minicpm2b_sft_bf16 (Done)
# python main.py --field general --model minicpm3_4b (Done)
# python main.py --field general --model minicpm_llama3_v_2_5 (Done)
# python main.py --field general --model chatglm3_6b (Done)
# python main.py --field general --model gpt4 (Done)
# python main.py --field general --model yi_chat_6b (Done)
# python main.py --field general --model yi_vl_34b (Done)
# python main.py --field general --model qwen_2_5_7b_instruct (Done)
# python main.py --field general --model internlm2_chat_7b (Done)
# python main.py --field general --model internlm2_5_chat_7b (Done)
# 2024.11.18
# python main.py --field general --model llama2_7b (Done)
# python main.py --field general --model llama2_13b (Done)
# python main.py --field rtl --model llama3_1_70b_instruct
# python main.py --field general --model llama3_8b_instruct (Done)
# python main.py --field general --model llama3_2_11b_vision_instruct (Done)
# python main.py --field general --model llama3_2_90b_vision_instruct (Done)
# python main.py --field general --model intern_chat_20b (Done)
# python main.py --field rtl --model internvl_mini_chat_2b_v1_5
# python main.py --field general --model internvl2_40b (Done)
# python main.py --field general --model internlm_xcomposer_vl_7b (Done)
# python main.py --field rtl --model internlm_xcomposer2_vl_7b
# python main.py --field general --model qwen_2_7b_instruct (Done)
# python main.py --field general --model qwen_vl_chat (Done)
# python main.py --field general --model qwen_vl (Done)
# python main.py --field general --model gpt4o (Done)
# python main.py --field general --model gpt4_turbo (Done)
# python main.py --field general --model gpt4v (Done)
# python main.py --field general --model qwen_2_72b_instruct (Done)
# 2024.11.16
# python main.py --field general --model gpt35 (Done)
# python main.py --field general --model qwen_2_0_5b_instruct (Done)
# python main.py --field general --model internvl2_8b (Done)
# python main.py --field rtl --model internvl_chat_v1_5
# python main.py --field general --model instructblip_flan_t5_xxl (Done)
# python main.py --field general --model llama3_1_8b_instruct (Done)
# python main.py --field general --model minicpm_v (Done)
# python main.py --field general --model minicpm_v2 (Done)
# python main.py --field general --model yi_vl_6b (Done)
- Blip-image-captioning-large
- GPT-3.5-turbo
- Llama2-7b-chat-hf
- Llama2-13b-chat-hf
- Llama3-8b-instruct
- Llama3.1-8b-instruct
- Llama3.1-70b-instruct
- mistrial-7b
- ChatGLM3 6b
- QWen2.5 7b Instruct
- QWen2.5 72b Instruct
- QWen2 7b Instruct
- QWen2 72b Instruct
- QWen2 0.5b Instruct
- Internlm chat 7b
- Internlm chat 20b
- Internlm2 chat 7b
- Internlm2 chat 20b
- Internlm2.5 7b chat
- Internlm2.5 20b chat
- miniCPM3-4b
- miniCPM-1b-sft-bf16
- miniCPM-2b-sft-bf16
- GPT-4o
- MiniGPT4-vicuna-7b
- MiniGPT4-vicuna-13b
- GPT-4
- GPT-4-turbo
- GPT-4-vision
- Reka series
- instructBlip flan T5 xl
- instructBlip flan T5 xxl
- blip2 flan t5 xl
- blip2 flan t5 xxl
- Llama3.2-11b-vision-instruct
- Llama3.2-90b-vision-instruct
- Kosmos2
- QWen VL
- QWen VL Plus
- QWen VL Max
- Bunny 1.0 3b (some bugs on the same device requirement)
- hpt 1.5 edge (some bugs about input of size)
- internlm xcomposer vl 7b
- miniCPM-V
- miniCPM-V2
- miniCPM-LLaMa3-V-2.5
- internvl2-8b
- internvl2-40b
- internvl-chat-v1.5
- mini-internvl-chat-2b-v1.5
- yi vl 6b
- yi vl 34b
- Gemini series
- Claude series
The data format in the project includes a JSON file and an image directory. The data for each subdomain follows the structure below:
- Image Directory: Such as
general/, containing image files. - JSON File: Such as
general.json, containing information like questions and answers.
An example structure of the JSON file is as follows:
{
"item_id": {
"statement": "Initial statement of the question",
"image": ["List of image names"],
"question": ["List of questions"],
"question_type": ["Type of questions"],
"answer": ["Answers to the questions"],
"explanation": ["Explanation of the answers"],
"modality": ["Data modality of the questions"],
"difficulty": ["Difficulty level of the questions"],
"ability": ["Model's testing ability on the questions"],
"source": "Source of the material"
}
}- General Knowledge
- Spec
- RTL
- Netlist
- Architecture
- Backend
Ensure Python 3.10 is installed, along with the necessary dependencies for the project.
An example is provided in main.py, demonstrating how to load data and benchmark models:
python:multimodalEDABenchmarking/main.py
startLine: 4
endLine: 33
The data loading functionality is implemented in utils/dataloader.py, supporting both single and batch data loading.
The benchmarking functionality is implemented in utils/parser.py, supporting testing of both single and batch questions.
python:multimodalEDABenchmarking/utils/parser.py startLine: 50 endLine: 140
Contributions to this project are welcome. Please submit a Pull Request or Issue to help us improve.
This project is licensed under the MIT License.