Skip to content

About

MM Benchmarking EDA

Resources

Stars

2 stars

Watchers

1 watching

Forks

Latest commit

 

History

110 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Multimodal EDA Benchmarking

简体中文CN

Updating

  • [√] supporting vLLM for text inference
  • [√] supporting vLLM for text & image inference

Project Overview

The Multimodal EDA Benchmarking project aims to provide a framework for evaluating and comparing the performance of different models on multimodal data. This project supports various data formats and models, helping researchers and developers better understand the capabilities of models in handling multimodal data.

Final Model List:

GPT
    3.5-Turbo  (text-only)
    4 (base64)
    4-Turbo (base64)
    4v (base64)
    4o (base64)
Reka
    Core (base64)
    Flash (base64)
    Edge (base64)
Gemini
    1.5 Pro (text-only)
Claude3.5
    Sonnet (text-only)
MiniGPT4
    Vicuna-7B (image-encoder)
QWen
    2-0.5B-Instruct (text-only)
    2-7B-Instruct (text-only)
    2-72B-Instruct (text-only)
    2.5-7B-Instruct (text-only)
    VL (image-encoder)
    VL-Chat (image-encoder)
InternLM
    Chat-20B (text-only)
    2-Chat-7B (text-only)
    2.5-Chat-7B (text-only)
    XComposer-VL-7B (image-encoder)
InternVL
    2-8B (image-encoder)
    2-40B (image-encoder)
InstructBLIP
    Flan-T5-XL (image-encoder)
    Flan-T5-XXL (image-encoder)
Blip2
    Flan-T5-XL (image-encoder)
    Flan-T5-XXL (image-encoder)
LlaMa
    2-7B-Chat-HF (text-only)
    2-13B-Chat-HF (text-only)
    3-8B-Instruct (text-only)
    3.1-8B-Instruct (text-only)
    3.2-11B-Vision-Instruct (image-encoder)
    3.2-90B-Vision-Instruct (image-encoder)
MiniCPM
    1B-SFT (text-only)
    2B-SFT (text-only)
    3-4B (text-only)
    V (image-encoder)
    V2 (image-encoder)
    LlaMa-3-V2.5 (image-encoder)
YiVL
    6B (image-encoder)
    34B (image-encoder)
YiChat
    6B (text-only)
ChatGLM3
    6B (text-only)
Kosmos
    2 (image-encoder)

Tested Models

2024.11.19

  • reka_flash
  • reka_edge
  • blip2_flan_t5_xl
  • blip2_flan_t5_xxl
  • instructblip_flan_t5_xl
  • kosmos2_patch14_224
  • minicpm2b_sft_bf16
  • minicpm3_4b
  • minicpm_llama3_v_2_5
  • chatglm3_6b
  • gpt4
  • yi_chat_6b
  • yi_vl_34b

2024.11.18

  • llama3_8b_instruct
  • llama3_1_70b_instruct (no general)
  • llama3_2_11b_vision_instruct
  • intern_chat_20b
  • internvl_mini_chat_2b_v1_5 (no general)
  • internvl2_40b
  • internlm_xcomposer_vl_7b
  • qwen_2_7b_instruct
  • gpt4o
  • gpt4_turbo
  • gpt4v

2024.11.17

  • MiniGPT4-vicuna7b (no general)
  • Llama 2-7b
  • Llama 2-13b
  • Internlm chat 7b
  • InternVL-Chat-V1.5
  • InternLM2 chat 7b
  • InternLM2 Chat 20b
  • InternLM2.5 Chat 7b
  • InternLM2.5 Chat 20b
  • Qwen2 72b instruct
  • QWen 2.5 7b instruct
  • QWen 2.5 72b instruct
  • MiniCPM-V

2024.11.15

  • GPT-3.5-turbo
  • QWen2 0.5b Instruct
  • internvl2-8b
  • instructBlip flan T5 xxl
  • Llama3.1-8b-instruct
  • miniCPM-V2
  • yi vl 6b

Running scripts

# 2024.11.19
# python main.py --field rtl --model reka_core
# python main.py --field general --model reka_flash (Done)
# python main.py --field general --model reka_edge (Done)
# python main.py --field general --model blip2_flan_t5_xl (Done)
# python main.py --field general --model blip2_flan_t5_xxl (Done)
# python main.py --field general --model gemini_1_5_pro
# python main.py --field general --model claude_3_sonnet
# python main.py --field general --model instructblip_flan_t5_xl (Done)
# python main.py --field general --model kosmos2_patch14_224 (Done)
# python main.py --field general --model minicpm1b_sft_bf16
# python main.py --field general --model minicpm2b_sft_bf16 (Done)
# python main.py --field general --model minicpm3_4b (Done)
# python main.py --field general --model minicpm_llama3_v_2_5 (Done)
# python main.py --field general --model chatglm3_6b (Done)
# python main.py --field general --model gpt4 (Done)
# python main.py --field general --model yi_chat_6b (Done)
# python main.py --field general --model yi_vl_34b (Done)
# python main.py --field general --model qwen_2_5_7b_instruct (Done)
# python main.py --field general --model internlm2_chat_7b (Done)
# python main.py --field general --model internlm2_5_chat_7b (Done)


# 2024.11.18
# python main.py --field general --model llama2_7b (Done)
# python main.py --field general --model llama2_13b (Done)
# python main.py --field rtl --model llama3_1_70b_instruct
# python main.py --field general --model llama3_8b_instruct (Done)
# python main.py --field general --model llama3_2_11b_vision_instruct (Done)
# python main.py --field general --model llama3_2_90b_vision_instruct (Done)
# python main.py --field general --model intern_chat_20b (Done)
# python main.py --field rtl --model internvl_mini_chat_2b_v1_5
# python main.py --field general --model internvl2_40b (Done)
# python main.py --field general --model internlm_xcomposer_vl_7b (Done)
# python main.py --field rtl --model internlm_xcomposer2_vl_7b
# python main.py --field general --model qwen_2_7b_instruct (Done)
# python main.py --field general --model qwen_vl_chat (Done)
# python main.py --field general --model qwen_vl (Done)
# python main.py --field general --model gpt4o (Done)
# python main.py --field general --model gpt4_turbo (Done)
# python main.py --field general --model gpt4v (Done)
# python main.py --field general --model qwen_2_72b_instruct (Done)

# 2024.11.16
# python main.py --field general --model gpt35 (Done)
# python main.py --field general --model qwen_2_0_5b_instruct (Done)
# python main.py --field general --model internvl2_8b (Done)
# python main.py --field rtl --model internvl_chat_v1_5
# python main.py --field general --model instructblip_flan_t5_xxl (Done)
# python main.py --field general --model llama3_1_8b_instruct (Done)
# python main.py --field general --model minicpm_v (Done)
# python main.py --field general --model minicpm_v2 (Done)
# python main.py --field general --model yi_vl_6b (Done)

TODO

Image Captioning

  • Blip-image-captioning-large

Text-only Model

  • GPT-3.5-turbo
  • Llama2-7b-chat-hf
  • Llama2-13b-chat-hf
  • Llama3-8b-instruct
  • Llama3.1-8b-instruct
  • Llama3.1-70b-instruct
  • mistrial-7b
  • ChatGLM3 6b
  • QWen2.5 7b Instruct
  • QWen2.5 72b Instruct
  • QWen2 7b Instruct
  • QWen2 72b Instruct
  • QWen2 0.5b Instruct
  • Internlm chat 7b
  • Internlm chat 20b
  • Internlm2 chat 7b
  • Internlm2 chat 20b
  • Internlm2.5 7b chat
  • Internlm2.5 20b chat
  • miniCPM3-4b
  • miniCPM-1b-sft-bf16
  • miniCPM-2b-sft-bf16

Multi-modal Model

  • GPT-4o
  • MiniGPT4-vicuna-7b
  • MiniGPT4-vicuna-13b
  • GPT-4
  • GPT-4-turbo
  • GPT-4-vision
  • Reka series
  • instructBlip flan T5 xl
  • instructBlip flan T5 xxl
  • blip2 flan t5 xl
  • blip2 flan t5 xxl
  • Llama3.2-11b-vision-instruct
  • Llama3.2-90b-vision-instruct
  • Kosmos2
  • QWen VL
  • QWen VL Plus
  • QWen VL Max
  • Bunny 1.0 3b (some bugs on the same device requirement)
  • hpt 1.5 edge (some bugs about input of size)
  • internlm xcomposer vl 7b
  • miniCPM-V
  • miniCPM-V2
  • miniCPM-LLaMa3-V-2.5
  • internvl2-8b
  • internvl2-40b
  • internvl-chat-v1.5
  • mini-internvl-chat-2b-v1.5
  • yi vl 6b
  • yi vl 34b

Multi-modal Model (image not ready)

  • Gemini series
  • Claude series

Data Format

The data format in the project includes a JSON file and an image directory. The data for each subdomain follows the structure below:

  • Image Directory: Such as general/, containing image files.
  • JSON File: Such as general.json, containing information like questions and answers.

An example structure of the JSON file is as follows:

{
    "item_id": {
        "statement": "Initial statement of the question",
        "image": ["List of image names"],
        "question": ["List of questions"],
        "question_type": ["Type of questions"],
        "answer": ["Answers to the questions"],
        "explanation": ["Explanation of the answers"],
        "modality": ["Data modality of the questions"],
        "difficulty": ["Difficulty level of the questions"],
        "ability": ["Model's testing ability on the questions"],
        "source": "Source of the material"
    }
}

Supported Subdomains

  • General Knowledge
  • Spec
  • RTL
  • Netlist
  • Architecture
  • Backend

Usage Instructions

Environment Setup

Ensure Python 3.10 is installed, along with the necessary dependencies for the project.

Running Example

An example is provided in main.py, demonstrating how to load data and benchmark models: python:multimodalEDABenchmarking/main.py startLine: 4 endLine: 33

Data Loading

The data loading functionality is implemented in utils/dataloader.py, supporting both single and batch data loading.

Benchmarking

The benchmarking functionality is implemented in utils/parser.py, supporting testing of both single and batch questions.

python:multimodalEDABenchmarking/utils/parser.py startLine: 50 endLine: 140

Contribution

Contributions to this project are welcome. Please submit a Pull Request or Issue to help us improve.

License

This project is licensed under the MIT License.

About

MM Benchmarking EDA

Resources

Stars

2 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages