← all notes
№ 004 2026-01-31 topic · Words in, starts out

How many stars can a model count?

On: Yelp star classifier — view on GitHub

Can a model tell you how many starts a yelp review got?

Turns out: mostly, yeah, as long as you’re not asking it to split hairs.

The task

I spent a week on a fun little NLP project: given the text of a Yelp restaurant review, predict how many stars (1 to 5) the person gave. No rating, no metadata, just the words. Here’s how it went.

The setup

I had about 35,000 reviews of New York City restaurants, already split into training, validation, and test sets. Each review came with a star rating from 1 (miserable) to 5 (loved it). My job was to predict that rating from the text alone, a five way classification problem, and the scoreboard metric was macro-F1, which just means every star level counts equally no matter how common it is.

Starting simple (on purpose)

It’s tempting to jump straight to a neural network, but the smart move is to build a dumb but solid baseline and see if anything can beat it.

So I started with a classic combo: TF-IDF + Logistic Regression. In plain terms, TF-IDF turns each review into a big list of numbers that reflect which words and word pairs show up, weighted so that distinctive words (“disgusting”, “phenomenal”) matter more than filler (“the”, “and”). Logistic Regression then learns which of those words lean toward which rating.

I tuned it a bit, trying single words vs. word pairs, how aggressively to drop rare words, how much to regularize, and landed at 0.62 macro-F1. Not bad for something that can be quicly trained.

Then I tried to beat it.

This is where it got humbling. I threw a few other things at the problem:

A linear SVM on the same features, basically tied, slightly worse. Character n-grams (looking at chunks of letters instead of whole words), worse. A small neural network in PyTorch that learns its own word representations and averages them together, landed at 0.60, right below the baseline.

The neural net was the interesting failure. With only 28,000 reviews and embeddings starting from scratch, it just didn’t have enough to learn from, it started memorizing the training data after about five passes while its validation score slid backwards. A well tuned bag of words quietly beat it.

Sometimes the boring model wins, and that’s a legitimate result worth reporting.

Findings

Aside from the score, an interesting part was where the model got things wrong. Breaking performance down by star rating:

1★ 2★ 3★ 4★ 5★

😃 0.72 😐 0.52 😐 0.56 😐 0.56 😃 0.71

The extremes are easy. A 1-star review screams (“never coming back”), a 5-star review gushes (“best meal of my life”), and the model picks up on that instantly.

The middle is a mess. A 3-star review often praises the food and complains about the wait in the same breath, so its vocabulary overlaps with both its neighbors. Almost every mistake the model made was off by one, calling a 3 a 4, or a 2 a 3. Which is… kind of what a human would do too. Is a “pretty good, but” review a 3 or a 4? Depends on the person.

Next steps

If I picked this back up, the obvious upgrade is giving the model real language knowledge instead of making it learn everything from 28k reviews, pretrained word embeddings, or fine tuning a small transformer like DistilBERT. I’d also try a loss function that knows the labels are ordered, so predicting a 5 for a 1-star review gets punished harder than predicting a 2. That directly targets the off-by-one problem.

Conclusion

The modeling was fun, but I spent just as much effort turning a messy experimentation notebook into a clean, reusable project: a proper Python package, a command line tool so anyone can retrain or predict with one line, automated tests, and a written report.

Disclaimer: the original task came from a lab in my Master of Data Science program. Everything past the initial notebook, the packaging, tests, CLI, and write up was a continuation of a fun project exploration.