Mrope td - #25
Mrope td#25danielezer wants to merge 5 commits into
Conversation
| t_mask = cos_offsets < mrope_section_t | ||
| h_mask = (t_end <= cos_offsets) & (cos_offsets < h_end) | ||
| w_mask = (h_end <= cos_offsets) & (cos_offsets < half_rd) | ||
|
|
There was a problem hiding this comment.
did you try to lower this with the triton fork to ktir?, from the CI logs, it seems to be only doing rms norm
There was a problem hiding this comment.
no, we didn't add ktir-cpu tests to the ci because it needs a gpu to run the vllm reference. Do you think this would be problematic for the compiler?
There was a problem hiding this comment.
im not 100% sure, but I have a feeling we might be missing some ops
There was a problem hiding this comment.
I tested the compilation + ktir-cpu locally and made some change to the kernel. Looks like the compilation works (a ktir file is generated), but the ktir-cpu simulation doesn't. It looks like the simulator doesn't have a proper lowering path for tl.arange, which causes the following error:
op = %cos_offsets_16 = tt.make_range()
.venv/lib/python3.13/site-packages/ktir_cpu/interpreter.py, _execute_op:
handler = dispatch(op.op_type)
if handler:
result = handler(op, context, self._env)
else:
> raise ValueError(f"Unknown operation: {op.op_type}")
E ValueError: Unknown operation: tt.make_range
Error executing tt.make_range on core 0: Unknown operation: tt.make_range
There was a problem hiding this comment.
Also, about the rms_norm being the only kernel tested with ktir-cpu: since we don't yet have a gpu to run the vllm reference, we don't execute the ktir tests yet. The rms_norm is now using a reimplemented reference in numpy, which can run on the cpu. This was done just so that we can't start a poc for the ci
Add a tensor descriptor kernel for mrope
Update td-test skill to avoid wrongly using tol=0 in tests