Repository navigation
Conversation
4edca4c to
5b640e7
Compare
4ba5dd5 to
51d4d55
Compare
|
|
||
| [This section is informational.] | ||
|
|
||
| The symbolic mapping of "lds-dma" scope affects how the _mutually inclusive_ |
There was a problem hiding this comment.
Is the term "mutually inclusive" used somewhere? It looks as if it was a fixed term, but the memory model only talks about "inclusive scopes"
| `Y` initiated from the same workgroup instance. On a target that performs DMA | ||
| operations at "cluster" scope, `Y` does not belong to any "workgroup" instance. | ||
| Thus `X` and `Y` do not have inclusive scope on this target even though they are | ||
| both associated with the same workgroup instance. If the same program is |
There was a problem hiding this comment.
What does "associated with" here mean, is that intentionally left vague?
| "workgroup". | ||
|
|
||
| Consider an operation `X` that specifies "workgroup" scope, and a DMA operation | ||
| `Y` initiated from the same workgroup instance. On a target that performs DMA |
There was a problem hiding this comment.
Should every occurrence of "DMA" in this paragraph be "LDS DMA"? (and maybe later occurrences as well?)
| call @llvm.amdgcn.asyncmark() | ||
| call @llvm.amdgcn.wait.asyncmark(0) | ||
| %val = call @llvm.amdgcn.av.load.b128(%global, "lds-dma") ; <-- | ||
| ``` |
There was a problem hiding this comment.
Maybe a note stating that that's not necessary for the LDS side would be helpful?
|
|
||
| A similar pattern is required with a DMA operation that writes to global memory. | ||
|
|
||
| ```llvm |
There was a problem hiding this comment.
Interesting point: if that thread writes something to %global (program-ordered-) before the global.store.async.from.lds call, you'd also need to make that available to lds-dma scope to establish a location-order that rules out a data race between the previous write and the store part of the async operation. Or should that be implicit in the global.store.async.from.lds operation?
| %val = load ptr addrspace(1) %global | ||
| ``` | ||
|
|
||
| Note how the acquire fence is used to establish visibility; it's ability to |
There was a problem hiding this comment.
| Note how the acquire fence is used to establish visibility; it's ability to | |
| Note how the acquire fence is used to establish visibility; its ability to |
|
|
||
| The AMDGPU backend further refines the LLVM scopes with the following | ||
| target-defined scopes and constraints: | ||
| The AMDGPU target defines the following LLVM scopes: |
There was a problem hiding this comment.
| The AMDGPU target defines the following LLVM scopes: | |
| The AMDGPU target defines the following scopes: |
| - If `S1` is a subscope of `S2`, then every instance `I1` of `S1` is a subset of | ||
| some instance `I2` of `S2`. `I1` is said to be a *subscope instance* of `I2`. | ||
| - If two scope instances `I1` and `I2` intersect, then their intersection is | ||
| the smaller of `I1` and `I2`. |
There was a problem hiding this comment.
I like that these are now presented as general properties of scopes.
| to*, as well as every scope instance `I2` such that `I1` is a *subscope | ||
| instance* of `I2`. |
There was a problem hiding this comment.
The part about subscope instances is redundant with the subscope instance definition above, right?
| This affects how operations are related in *inclusive scopes* when determining | ||
| availability and visibility. For example, {ref}`DMA | ||
| operations<amdgpu-dma-memmodel>` require {ref}`explicit availability and | ||
| visibility<amdgpu-dma-visibility>` operations at scopes below the corresponding |
There was a problem hiding this comment.
Why "below"? Depending on how you interpret that sentence, "above" would make more sense. But maybe "...at the corresponding "lds-dma" scope." would be clearer?
Part of a stack: