|
cpu.o
|
make contiguous cuda
|
2024-05-23 02:27:30 -03:00 |
|
cuda.cu.o
|
broadcast allreduce nccl
|
2024-05-24 18:19:39 -03:00 |
|
distributed.o
|
add argument src to broadcast distributed
|
2024-05-25 09:38:42 -03:00 |
|
Makefile
|
fix send to other devices and compile on cloud
|
2024-05-29 17:48:37 -03:00 |
|
tensor.o
|
implemented mpi for distributed run
|
2024-05-24 15:12:53 -03:00 |