Volt says you can drop the convolutions. On a small 3D dataset, the bias earns its keep.
Volt argues that a vanilla transformer, with the right recipe and enough data, no longer needs a convolution-specialized backbone for 3D scene understanding. We took that seriou...
3dtransformerconvolution