Chirag Jain
|
50896ec574
Make nvcc threads configurable via environment variable (#885)
|
il y a 9 mois |
Tri Dao
|
393882bc08
[LayerNorm] Implement LN with parallel residual, support dim 8k
|
il y a 1 an |
Tri Dao
|
dc08ea1c33
Support H100 for other CUDA extensions
|
il y a 1 an |
Tri Dao
|
8c6609ae1a
[LayerNorm] Support all dimensions up to 6k (if divisible by 8)
|
il y a 2 ans |
Tri Dao
|
39ed597b28
[LayerNorm] Compile for both sm70 and sm80
|
il y a 2 ans |
Tri Dao
|
fa6d1ce44f
Add fused_dense and dropout_add_layernorm CUDA extensions
|
il y a 2 ans |