@assaf No consumer hardware exists to run 70b models locally without spending as much as a car or two. Everyone talking about how cool it is to run a model locally overlooks the hard to measure (model quality) and how important tokens per second is.
Post
March 9, 2024