Existing #SMBC comics data compiled into convenient datasets
I'm happy with how this turned out so far.
TODO
- Another round of OCR (anything better than tesseract?)
- LLM OCR (Would that be expensive?)
- Merge the ~4 datasets into one for convenience
https://github.com/matthewdeanmartin/smbc_scraper?tab=readme-ov-file#smbc_scraper