Conversation
A load unit cannot always splat a single element: when its smallest access granularity is wider than the source, the hardware loads the whole granule and splats that, so one element per sub-SIMD group ends up live rather than one for the whole vector. Give the IR the two things a consumer needs to finish that. ktdf.read_from_fifo gains an optional `splat` attribute, #ktdf.splat<...>, whose one mode -- first_subsimd_lane_to_all_subsimd_lanes -- names the shuffle that closes the gap. It names the shuffle rather than asserting that something is owed, so a reader learns what to build and not only that it must build. It is inherent, not discardable, because it changes what the op means: a read carrying it hands out a value the shuffle has not been applied to yet. The attribute survives onto a memref-typed re-read of the same slot, since reading a buffer performs no shuffle. The mode carries no width. A sub-SIMD group is as wide as the compute unit says, so feature::SIMD gains getSubSimdLanes, mirroring getLanes, over the sub_simd_lanes map device files already declare and nothing read. Stating the width in a device pattern instead would make two places have to agree about one arch fact. A trailing-optional attribute has no default in the generated builder, so read_from_fifo also gains the two-argument builder every existing caller uses: a read that owes no shuffle should not have to spell an absent mode. Signed-off-by: Masoud Ataei <Masoud.Ataei.Jaliseh@ibm.com>
Contributor
|
If we're going forward with this, I'd like to note that the new attribute does not have matching changes in the corresponding feature test function for |
feature::SIMD::test() silently ignored sub_simd_lanes requirements.
Add the provider->=required check mirroring the lanes block, and add a
matching SUBCASE("sub_simd_lanes") in the Features unit test.
Signed-off-by: Masoud Ataei <Masoud.Ataei.Jaliseh@ibm.com>
Signed-off-by: Masoud Ataei <Masoud.Ataei.Jaliseh@ibm.com>
msdataei
marked this pull request as ready for review
September 24, 2026 13:39
msdataei
requested review from
Prasanth-Chatarasi,
acgatea1,
ani300,
bmahjour,
lupalby,
mudhakar,
viji560 and
vswagath1989
as code owners
September 24, 2026 13:39
bmahjour
requested changes
Sep 24, 2026
ktdf.read_from_fifo
Signed-off-by: Masoud Ataei <Masoud.Ataei.Jaliseh@ibm.com>
Signed-off-by: Masoud Ataei <Masoud.Ataei.Jaliseh@ibm.com>
msdataei
force-pushed
the
msd_spalt_fp32
branch
from
September 24, 2026 21:17
d4fbec9 to
4d97a6d
Compare
Contributor
Author
|
the changes in this PR is used in the PR torch-spyre/dataflow-scheduler#170 |
Contributor
Author
|
Upstream PR: #4697 |
KFAFSP
reviewed
Sep 28, 2026
| static constexpr StringLiteral kSplatAttrName = "splat"; | ||
| static constexpr StringLiteral kZeroPadAttrName = "zero_pad"; | ||
| static constexpr StringLiteral kLanesAttrName = "lanes"; | ||
| static constexpr StringLiteral kSubSimdLanesAttrName = "sub_simd_lanes"; |
Contributor
There was a problem hiding this comment.
Keeping with the style of the existing attribute, this should be called sub_lanes I think.
There was also the open design question on whether SIMD lanes could be multidimensional, i.e., f16 = 64x64 to indicate matrix-type accelerators. I don't think this is quite the way to go, but it would be good to know whether sub SIMD lanes would be any different in that regard.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
When a FIFO delivers fewer than a full SIMD width of elements, a consumer
needs two pieces of information to form the final vector: that a shuffle is
required, and which shuffle to apply. This change encodes both in the IR via a
new optional
splatattribute onktdf.read_from_fifoand a newgetSubSimdLanesaccessor onfeature::SIMD.ktdf.read_from_fifo— new optionalsplatattributeThe op gains
OptionalAttr<KTDF_SplatModeAttr>:$splat. Its one mode,first_subsimd_lane_to_each_subsimd, names the shuffle that fills everylane of each sub-SIMD group with the value held by its first lane. Naming the
shuffle rather than asserting a bare obligation means a reader learns both that
a shuffle is owed and how to build it.
Because a trailing
OptionalAttrproduces no default in the generated builder,a two-argument builder is added for the common case —
(result, fifo_slot)—so existing callers do not have to spell an absent mode.
feature::SIMD— newgetSubSimdLanesaccessorThe shuffle mode refers to sub-SIMD group width, but the width is an arch fact,
not part of the attribute.
feature::SIMDgainsgetSubSimdLanes()andgetSubSimdLanes(Type), mirroring the existinggetLanespair, over thesub_simd_lanesmap that device files already declare. The verify and testpaths are extended to cover the new key, and the
intrinsics.mlirfixture isupdated to declare
sub_simd_laneson its test device. Keeping the width inthe arch avoids two places having to agree about one fact.
Tests
read-from-fifo-op.mlir— round-trip checks forsplaton both tensorand memref result forms.
features-invalid.mlir— verifier rejects a non-map value forsub_simd_lanes.Features.cpp—SUBCASE("sub_simd_lanes")covers the test predicate:absent provider satisfies empty requirement; absent provider fails non-empty
requirement; narrow provider fails wider requirement; wider provider satisfies
narrower requirement.