david/flash-attention

Autor	SHA1 Mensaje	Fecha
Zhihao Shen	30e1ef0f79 minify torch.torch.int32 to torch.int32 (#1237)	hace 2 meses
Ying Zhang	cdbbe844b1 minor changes to unpad_input test util func	hace 3 meses
Tri Dao	abbc131173 [LayerNorm] Switch from CUDA to Triton implementation	hace 11 meses
Kevin Hu	07005806ff Add BigCode converters (#532)	hace 1 año
Kevin Hu	4c91621a5e Inverse state dict for BERT (#527)	hace 1 año
Tri Dao	f1a73d0740 Run isort and black on python files	hace 1 año
Kiarash Jamali	684196b8c5 Allow rotary embeddings for Bert (#363)	hace 1 año
Tri Dao	96d10f6545 Implement LLaMa	hace 1 año
Tri Dao	88173a1aaf [FusedDense] Support relu, rename FusedDenseGeluDense -> FusedMLP	hace 1 año
Tri Dao	ff34123bd4 Reorder LN in Block, support OPT	hace 1 año
Tri Dao	714c1b4f0f [Bert] Fix embedding layer norm before embedding dropout	hace 1 año
Tri Dao	c6ecd40a59 Tweak CrossEntropyLoss to take process_group in init	hace 2 años
Tri Dao	dff68c2b22 Add smoothing for CrossEntropyParallel, rename to CrossEntropyLoss	hace 2 años
Tri Dao	e68ebbe89a Simplify FusedDense	hace 2 años
Tri Dao	13cdceb377 Implement last_layer_subset optimization for BERT	hace 2 años
Tri Dao	5fb6df0e04 Implement BERT	hace 2 años