Recently I’ve been looking into (and contributing a little bit to) river, an online machine learning library for Python.

Even though “standard” (i.e offline) machine learning has proved to be powerful, I think that the ability to handle data streams (and all the nasty things that go with it, such as unknown number of classes, concept drift, etc.) is very much underestimated today, and therefore understudied.

But what can I do about this? Diving deep enough into online ML to make a discovery feels out of reach (for me at least).

However a non-negligible part of the success of offline machine learning is due to the abundance of challenges and benchmarks, and this is something that is currently missing in online ML (this is not a 100% original thought: see this blog post from Max Halford, the main character behind river).

So I figured I’d build one! It’s called fontaine and it is a font recognition challenge: you get a stream of textbox crops, i.e a few words written on top of some background image (typically what would come out of an OCR phase), and your task is to find which font was used to draw the text.

Here is what I like about this challenge:

So basically whenever the challenge feels to easy, we just update the configuration file to make it harder (mwahahaha).

The leaderboard is currently occupied by me, myself and I: I created a little CNN-based model called GriftNet, with one frozen layer of ResNet-18 and a growable classification head. Feel free to annihilate it with much better models, that’s the whole point.

Hopefully fontaine gets enough interest and we start seeing some new models and patterns emerge :)