大模型训练技术论文

A Reading List for MLSys

An Overview of Distributed Methods | Papers With Code

ZeRO: Memory Optimizations Toward Training Trillion Parameter Models

https://arxiv.org/abs/1910.02054

Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

https://arxiv.org/abs/2104.04473

Reducing Activation Recomputation in Large Transformer Models

https://arxiv.org/abs/2205.05198

Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

https://arxiv.org/abs/1909.08053

Fully Sharded Data Parallel: faster AI training with fewer GPUs

Fully Sharded Data Parallel: faster AI training with fewer GPUs Engineering at Meta -

GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding

https://arxiv.org/pdf/2006.16668.pdf

GSPMD: General and Scalable Parallelization for ML Computation Graphs

https://arxiv.org/pdf/2105.04663.pdf

Automatic Cross-Replica Sharding of Weight Update in Data-Parallel Training

https://arxiv.org/abs/2004.13336v1

评论
添加红包

请填写红包祝福语或标题

红包个数最小为10个

红包金额最低5元

当前余额3.43前往充值 >
需支付:10.00
成就一亿技术人!
领取后你会自动成为博主和红包主的粉丝 规则
hope_wisdom
发出的红包

打赏作者

张博208

你的鼓励将是我创作的最大动力

¥1 ¥2 ¥4 ¥6 ¥10 ¥20
扫码支付:¥1
获取中
扫码支付

您的余额不足,请更换扫码支付或充值

打赏作者

实付
使用余额支付
点击重新获取
扫码支付
钱包余额 0

抵扣说明:

1.余额是钱包充值的虚拟货币,按照1:1的比例进行支付金额的抵扣。
2.余额无法直接购买下载,可以购买VIP、付费专栏及课程。

余额充值