Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
58 commits
Select commit Hold shift + click to select a range
54b797c
refactor(runtime)!: adopt modern Infini stack
voltjia Jul 16, 2026
93a161f
docs: define Infini stack repository boundaries
voltjia Jul 16, 2026
da76f3f
refactor: own Infini stack integration build
voltjia Jul 16, 2026
05fad31
docs: document the InfiniLM stack build
voltjia Jul 16, 2026
b15909a
ci: validate the modern NVIDIA stack
voltjia Jul 16, 2026
2153f7d
fix: keep static graph cache metadata dynamic
voltjia Jul 24, 2026
1a8d3d4
refactor(ops): use canonical InfiniOps APIs
voltjia Aug 11, 2026
2d34490
fix(build): select required linked InfiniOps providers
voltjia Aug 11, 2026
d9aaaf4
fix(build): limit compiled models to Qwen3
voltjia Aug 11, 2026
f282577
feat(distributed): add eager point-to-point wrappers
voltjia Aug 11, 2026
fea3dad
feat(ops): restore prepacked linear execution
voltjia Aug 11, 2026
69d0ff2
fix(ops): match canonical Gemm call schema
voltjia Aug 11, 2026
4c541c9
fix(sampling): support scalar sample outputs
voltjia Aug 11, 2026
c84a862
fix(nn): pass RoPE output handles by value
voltjia Aug 11, 2026
9e380f7
fix(attention): include variable-length MHA declaration
voltjia Aug 11, 2026
30d90e8
fix(engine): use local static graph assertions
voltjia Aug 11, 2026
5662890
fix(config): reject unsupported Qwen3 linear bias
voltjia Aug 11, 2026
54d95b3
test: cover modern runtime contracts
voltjia Aug 11, 2026
87a4087
docs: update modern Qwen3 support boundary
voltjia Aug 11, 2026
93ae2c6
refactor(build): select InfiniOps implementations from JSON
voltjia Aug 13, 2026
2a1880b
style: format InfiniOps attention adapters
voltjia Aug 13, 2026
2809bec
fix(graph): update paged replay metadata through tensors
voltjia Aug 13, 2026
9b00e5f
fix(graph): fall back to eager paged decode under tensor parallelism
voltjia Aug 13, 2026
3b4543b
fix: restore modern InfiniCore API compatibility
voltjia Aug 18, 2026
a6f8c59
fix(runtime): sync post-migration InfiniCore updates
voltjia Aug 25, 2026
52c5073
style: format migrated sources
voltjia Aug 25, 2026
92d809f
fix(runtime): keep receive buffers mutable
voltjia Aug 25, 2026
2792bad
fix(build): exclude unsupported MXFP4 path
voltjia Aug 25, 2026
eeb55f0
fix(packaging): preload Torch shared libraries
voltjia Aug 25, 2026
8cc81f3
refactor(runtime): replace deprecated causal softmax backend
voltjia Aug 26, 2026
49a7f26
style: format causal softmax adapter
voltjia Aug 26, 2026
55e1622
feat(runtime): unlock validated InfiniOps capabilities
voltjia Aug 26, 2026
bc31217
fix(runtime): select available InfiniOps implementations
voltjia Aug 26, 2026
2b4e9c8
fix(runtime): enable available Moore InfiniOps paths
voltjia Aug 27, 2026
7e6d100
fix(runtime): accept native MetaX device name
voltjia Aug 27, 2026
6261df2
fix(core): validate Cat inputs against tensor rank
voltjia Aug 27, 2026
ebce722
feat(models): enable validated MiniCPM Eagle path
voltjia Aug 27, 2026
9445436
style(runtime): group ModelRunner imports
voltjia Aug 27, 2026
b700358
fix(graph): enable safe paged decode capture under TP
voltjia Aug 31, 2026
cf92e39
feat(moore): enable modern flash attention
voltjia Sep 2, 2026
fbcad61
style(moore): format attention integration
voltjia Sep 2, 2026
822cd66
build(moore): configure the modern stack
voltjia Sep 2, 2026
8a2f6e2
fix(distributed): avoid transient runtime teardown
voltjia Sep 2, 2026
56accff
feat(cambricon): enable tensor-parallel flash attention
baominghelly Sep 2, 2026
ddc27f1
feat(iluvatar): enable modern Infini stack inference (#557)
gongchensu Sep 8, 2026
a8b4c1e
feat(hygon): enable the modern Infini stack (#564)
gongchensu Sep 9, 2026
70b5154
feat(ascend): enable InfiniOps flash attention (#558)
baominghelly Sep 9, 2026
90ddf65
perf(nvidia): reduce refactored inference overhead (#565)
voltjia Sep 9, 2026
dcbf3b6
fix(ascend): use CommInitAll for single-node TP (#566)
baominghelly Sep 10, 2026
a84c58a
fix(ascend): keep unsafe operators out of device graphs (#568)
baominghelly Sep 11, 2026
1e261e5
fix(allocator): reuse pinned free blocks (#569)
voltjia Sep 11, 2026
9ef8a20
fix(runtime): address inference migration review feedback
voltjia Sep 21, 2026
6661765
fix(attention): remove paged attn
Sep 23, 2026
1b7f61f
fix(inference): preserve migration interfaces and backend selection
voltjia Sep 23, 2026
03dab96
fix-inference-keep-thead-as-canonical-ppu-device
wooway777 Sep 24, 2026
e649abc
fix(attention): select active flash attention provider
wooway777 Sep 27, 2026
d2a8504
fix(attention): remove legacy flash bridges
wooway777 Sep 27, 2026
f94d83d
style: format attention adapter
wooway777 Sep 27, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
74 changes: 74 additions & 0 deletions .github/ci_config_nvidia.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,74 @@
repo:
url: https://github.com/InfiniTensor/InfiniLM.git
branch: main

github:
status_context_prefix: "ci/infinilm"

platforms:
nvidia:
image:
dockerfile: images/nvidia/
build_args:
BASE_IMAGE: nvcr.io/nvidia/pytorch:25.12-py3
CUDA_ARCH: sm_80,sm_86,sm_89,sm_90
APT_MIRROR: https://mirrors.tuna.tsinghua.edu.cn/ubuntu
PIP_INDEX_URL: https://pypi.org/simple
InfiniCore_BRANCH: __Branch_Name__
docker_args:
- "--user=root"
- "--network=host"
- "--privileged"
- "--cap-add=ALL"
- "--pid=host"
- "--ipc=host"
- "--workdir=/workspace"
volumes:
- /data:/data
- /data-aisoft:/data-aisoft
- /data-aisoft/artifacts/CI_nvidia_test/__WORKSPACE__:/artifacts
setup: pip install .[dev] --no-build-isolation
jobs:
gpu_inferencetest:
type: inferencetest
resources:
ngpus: [1, 4]
gpu_style: nvidia
shm_size: 64g
timeout: 3600
stages:
- name: test
run: python InfiniLM/examples/test_infer.py --device nvidia --model=/data-aisoft/mechdancer/models/9g_8b_thinking/
gpu_benchtest:
type: benchtest
resources:
gpu_style: nvidia
shm_size: 64g
timeout: 3600
env:
TEST_PARAM: ['default']
stages:
- name: test
run: python InfiniLM/examples/bench.py --device nvidia --model=/data-aisoft/mechdancer/models/9g_8b_thinking/ --input-len=256,1024 --output-len=256,1024 --batch-size=8 <TEST_PARAM>
gpu_accuracytest:
type: accuracytest
resources:
gpu_style: nvidia
shm_size: 64g
timeout: 3600
env:
TEST_PARAM: ['--bench mmlu']
stages:
- name: test
run: python InfiniLM/test/bench/test_benchmark.py --device nvidia --model /data-aisoft/mechdancer/models/9g_8b_thinking/ --bench mmlu --backend cpp --max-new-tokens 5 --cache-dir /data-aisoft/pepe/datasets/ --split=val <TEST_PARAM>
gpu_servicetest:
type: servicetest
resources:
shm_size: 64g
env:
MODEL_LIST: 9g_8b_thinking
ENGINE: InfiniLM
TEST_PARAM: ['default']
stages:
- name: test
run: python InfiniLM/scripts/test_perf.py --verbose
6 changes: 3 additions & 3 deletions .github/workflows/ci_test.yml
Original file line number Diff line number Diff line change
Expand Up @@ -41,9 +41,9 @@ jobs:
ci:
if: github.event_name == 'workflow_dispatch'
needs: check-format
uses: InfiniTensor/ci/.github/workflows/infinilm-ci.yml@infiniCore_ci
uses: InfiniTensor/ci/.github/workflows/infinilm-ci.yml@refactor/adopt-modern-infini-stack
with:
config_path: .github/ci_config.yaml
ci_ref: infiniCore_ci
config_path: .github/ci_config_nvidia.yaml
ci_ref: refactor/adopt-modern-infini-stack
infinicore_branch: ${{ github.event.inputs.infinicore_branch || 'main' }}
secrets: inherit
1 change: 1 addition & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -10,6 +10,7 @@ python/infinilm/lib/*.so
.vscode/

*.sh
!test/native/run_tests.sh
model_weight/

# Python
Expand Down
24 changes: 23 additions & 1 deletion CONTRIBUTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -58,7 +58,29 @@ Existing branch names may use the legacy format `issue/###`, followed by a suffi

# Development Guide

Refer to [ReadMe](README.md) and [Adapt New Models](MODELS.md)
Refer to [ReadMe](README.md) and [Adapt New Models](MODELS.md).

Run the migrated stack-builder unit tests with:

```shell
python -m unittest discover -s test/scripts -p "test_build_infini_stack.py" -v
```

Check the Core-backed build commands without creating build output:

```shell
python scripts/build_infini_stack.py --infinicore-root ../InfiniCore --dry-run --jobs 1 --cuda-arch sm_80
```

Run the static migration contracts with:

```shell
python -m unittest discover -s test/static -p "test_*.py" -v
```

Native stack changes must be validated in the owning component repository
first. After validation, update the corresponding submodule pin in InfiniCore
before validating InfiniLM.

## Troubleshooting

Expand Down
Loading
Loading