mentioned in paper,"given the same batch size, GEAR significantly reduces the peak memory compared to FP16 baseline, increasing the maximum severing number (i.e., batch size) from 3 to 18"

Why does the peak memory of GEAR decrease at a slope that is 1/4 to 1/3 of FP16 as the batch size increases? Is it due to the quantization of activations?
mentioned in paper,"given the same batch size, GEAR significantly reduces the peak memory compared to FP16 baseline, increasing the maximum severing number (i.e., batch size) from 3 to 18"

Why does the peak memory of GEAR decrease at a slope that is 1/4 to 1/3 of FP16 as the batch size increases? Is it due to the quantization of activations?