← Back to all articles
Reddit r/MachineLearningSeptember 8, 2026

I reduced image-processing token usage by ~95% compared with GPT-4o direct vision, while maintaining roughly the same accuracy.How significant is that?[P]

Excerpt

I'm testing a new approach for reducing the cost of image-based LLM inference. I evaluated it on the MOMA Graph benchmark , using 1,315 questions . Compared with using GPT-4o to process the original images directly, I observed approximately: ~95% lower token usage roughly the same accuracy as the GPT-4o direct-image baseline I'm intentionally not sharing implementation details yet because the method is still under development. I'm mainly trying to understand how strong the result itself is. If t