The product promises depth: hundreds of exercises on one narrow grammar topic. Depth has an obvious failure mode. Forty exercises can easily feel like the same exercise forty times, and a learner who feels that will stop, usually without saying why.

Two questions follow. Are these exercises genuinely different? And even if they are, does the order a learner meets them in make them feel repetitive? A learner never sees a similarity matrix. A learner sees a sequence.

Embeddings answered the wrong question

The obvious first step was off-the-shelf sentence embeddings. They failed in two directions. Two exercises about travel that practise different grammar came out as very similar. Two exercises about different topics that use the same sentence frame and the same answer pattern came out as different. For a learner it’s exactly the other way round: the travel pair feels varied, and the second pair feels like a repeat.

So the similarity score had to describe what “repetitive” means for this product. It ended up with four parts, each normalised and weighted before comparing exercises:

  • Words in the filled-in text, which catches the same vocabulary field coming up again and again.
  • Character patterns of the correct answers, which catches the same endings drilled with different words around them.
  • Structure: gap count, text and answer length, where the gaps sit.
  • Part-of-speech patterns from spaCy, which catches the same sentence scaffold dressed in new vocabulary.

The exact weights matter less than the decision behind them: similarity here is a product definition, not a generic NLP quantity.

Making a score usable

A section with 40 exercises has 780 pairs, and a list of 780 numbers is not something anyone can act on. The admin area shows the same data at three zoom levels: a heatmap for where the overlap is dense, a dendrogram for whether it forms families of near-variants, and a side-by-side view to decide whether a close pair is truly redundant or just related.

Similarity heatmap for one section: each cell is the similarity between two exercises, coloured from low to high.
The heatmap shows where overlap is dense. Anything above about 0.5 is worth a direct look as a possible near-duplicate.
Two exercises side by side with their gaps and answers highlighted, at 32.92% similarity.
The pair view is where the decision happens: redundant, or merely related?

The metric that lied to me

To compare orderings I needed a single number per section. My first one divided the average similarity of neighbouring exercises by the average similarity of all pairs. Lower meant better spacing. It looked sensible and it moved in the right direction after reordering.

Then I added a batch of new exercises to a section without changing anything about the order, and the score changed anyway. The ratio depended on how many exercises there were and how similar the new ones were to everything else, not only on the order. A metric that improves when nothing has improved is worse than no metric, because you believe it.

The replacement asks a rank question instead: where do the neighbouring pairs sit among all pairs in the section, from least to most similar? If neighbours come from the dissimilar end, the order is good, whatever the section’s size. A second measure looks a few steps ahead, weighting the next exercise most and the fifth one least, because a learner feels the immediate repeat most.

Reordering without breaking anyone’s place

The ordering itself is a classic heuristic: start with a greedy sequence that always picks the least similar next exercise, then improve it with 2-opt swaps until no swap helps. It’s quick and it isn’t optimal, which is fine. The interesting part is the constraint.

Learners move through a section in order, and their progress is stored as the number of the last exercise they completed. Reorder the whole section and a returning learner might repeat something or skip something. So the rule is: anything at least one learner has completed stays where it is.

Exercises 1 to 5 have been completed by at least one learner and stay frozen in their old order. Exercises 6 to 12 have not been seen and are reordered. The first free exercise is chosen to be the least similar to the last frozen one, because that is where a returning learner continues.
Constraint first, optimisation second. A weaker order that respects every learner’s place beats a better one that breaks it.

One detail sits at the boundary. The first exercise after the frozen part is chosen to be as unlike the last frozen one as possible, because that is exactly where a returning learner continues. When a reorder does renumber exercises, every learner’s progress pointer is remapped in the same database transaction. Near-duplicates are not deleted, by the way. They are spaced apart. A close variant twenty exercises later is good repetition; the same thing twice in a row is not.

Similarity of each exercise to the next one to five exercises, before reordering, with high values such as 0.57 and 0.47 in the next-exercise column.
Before: the column for the very next exercise shows values such as 0.57 and 0.47, so neighbouring exercises overlap noticeably.
The same view after reordering, with most next-exercise values below 0.20.
After: most values in that column fall below 0.20.

If you sequence content for learners

  1. Define similarity from the user’s experience, not from a generic model’s idea of meaning.
  2. Test a metric by changing something irrelevant. If it moves, it’s lying.
  3. Treat existing users’ progress as a hard constraint and optimise inside it.
  4. Space near-duplicates apart instead of deleting them; repetition is good, back-to-back repetition isn’t.