I think there are many other significant inference improvements which an be built out that are being ignored because the API between client and inference stack would be tricky to nail down.
This is not a criticism of the research but instead the presentation but the repo looks very fishy, it isn't clear that this is from a bunch of researchers from Berkley. It also does not make very clear (on the GitHub) what optimizations or performance they are targeting.