How does a feed learn what you like?

I built a TikTok-style recommendation engine from scratch to answer a question I couldn't shake: does the app know a video is about cats, or does it just know I keep watching things like it? This page walks through the build, from a scoring rule you could run on paper to Google Gemini embeddings and a FAISS vector index.

Live at social-media-algo-code.onrender.com (a free Render instance, so the first load can take a moment to wake up). Code at github.com/ayushxpatne/social-media-algo-code.

A real run of the first version, ten interactions in

Only three topics exist: cats, dogs and space. A like scores +1, a dislike scores -1. The user saw 3 cat videos and liked 2, saw 5 dog videos and liked 4, saw 2 space videos and liked neither.

TopicShownLikedMean score
dogs540.60
cats320.33
space20-1.00

the dumb algorithmmean of plus-one, minus-one

preference_score = mean(scores per topic)
dogs: 0.60   cats: 0.33   space: -1.00

next batch: mostly dogs, some cats, no space

This is the whole system at the start of the project: three topics, one score per interaction, and a mean. It ranks the topics correctly, but it does not know why a video is a dog video. Someone had to label it that way by hand. Sections below replace that label with an embedding, and the mean with a vector average.

A scoring rule you could run by hand

The first version has no machine learning in it at all. Every video already carries a hand-assigned topic. Each interaction adds a score to that topic, and a preference score is just the mean score per topic so far. The hero example above is a real run of it.

To turn that ranking into a feed, the batch of videos to show is split across topics in proportion to preference score:

current_category_i = batch_size * preference_score_i / sum(all preference_scores)

With the scores above, a batch of 20 videos works out to about 13 dog videos and 7 cat videos, and zero space videos. That is exploitation: showing more of what a user already likes. Its opposite, exploration, means showing something new on purpose to learn whether the user likes it too. This first version does no exploration at all.

This rule works, but it does not scale. A real platform has thousands of categories, not three, and someone has to hand-label every video's topic before the mean can even be computed. Growing the interaction types helps a little:

interactions = {
    # view_time under 2s counts as a skip, under 7s as a short watch, over 7s as a long watch
    "view_time": {"skip": -1, "short": -0.5, "long": +1.5},
    "like": +1,
    "comment": +2,
    "share": +3,
    "save": +3,
}

More signals give a finer-grained score, but the core problem stays: the system still only knows a video's topic because a human typed it in. It has no way to look at a new video and place it anywhere.

Teaching the system what a video is

A machine learning model cannot read the word "cat". It only works with numbers. The fix is an embedding: a list of numbers that acts like a coordinate for a piece of content, placed so that similar things end up near each other.

Picture a toy space with three axes: how alive something looks, how cute or friendly it feels, and how mechanical it is. A dog might sit at (0.7, 0.9, 0.2), a cat close by at (0.8, 0.7, 0.3), and a car far off at (0.1, 0.2, 0.9). Nobody chose those numbers by hand. A neural network learns them by looking at huge amounts of text (or video, or audio) and gradually placing things that appear in similar contexts closer together.

alive to lifeless cute to mechanical dog (0.7, 0.9) cat (0.8, 0.7) car (0.1, 0.2) close together
Dog and cat land near each other; car sits far away. The axes here are made up to be readable. Google Gemini Embeddings, which the project actually uses, place every video in 3072 dimensions with no axis that a person can name, but the same idea holds: distance in that space stands in for similarity in meaning.

So does the algorithm know a video is a cat video? Yes and no. It knows the video's position in a 3072-number space where cat videos cluster together. It has no label that reads "cats". In this project the inputs are still text, categories and descriptions converted to embeddings, not the raw video frames, but the principle is the one real platforms use on actual pixels and audio.

Building one number for a whole user

To recommend content, the system also needs a single point that represents what one user likes: a user embedding. It is built as a weighted average of the embeddings of videos that user engaged with, weighted by the score each video got.

user_embedding = sum(embedding_i * score_i for each watched video i) / sum(all scores)

Keeping the feed fresh

A first working version of the FAISS-based feed had two problems. Videos arrived in clumps from the same topic, and after about 20 videos the feed would run out of similar items to show and stop. Two small changes fixed both.

A sliding window instead of a lifetime average

The user embedding was originally a weighted average of a user's entire watch history. That means someone who binged cat videos yesterday still gets a cat-heavy feed today even if they've since moved on. The fix is a sliding window: only the last 5 videos with a score of 2 or higher (meaning they were properly watched, not just skipped past) feed into the user embedding.

filtered_df = history_df[history_df["score"] >= 2].tail(5)
for row in filtered_df.itertuples():
    embedding_matrix.append(row.embeddings)
    weights.append(int(row.score) + 1)

weights = weights / sum(weights)
user_embedding = weighted_average(embedding_matrix, weights)

This is a stand-in for time decay, where older interactions matter less by a formula rather than a hard cutoff. The project does not record timestamps, so a fixed window of 5 recent, high-score videos is the simpler substitute: weight_i = score_i * e^(-lambda * time_elapsed_i) is the version that would replace it if timestamps were tracked.

The 70-30 mix

The running-out problem and the clumping problem share one fix. Instead of asking FAISS for 10 similar videos every batch, the system asks for only 7, then adds 3 videos picked at random from the whole database, and shuffles the combined 10 before sending them to the user. A video already shown gets swapped for a random one on the spot, so the feed never has fewer than 10 videos to show.

Try the mix yourself

Drag the slider to change how many of the next 10 videos come from the user's taste (FAISS neighbours of the user embedding, mint) versus picked at random for exploration (apricot). The project's own default is 7 similar to 3 random.


from your taste
7
FAISS neighbours of the user embedding
random exploration
3
picked from the whole database

Why there is no separate ranking step

Every earlier version of this project sorted videos by preference score before serving them, a separate ranking step. Once FAISS entered the picture, that step quietly disappeared, because a similarity search already returns results ordered by distance: the closest video to the user embedding comes first. Retrieval and ranking collapsed into the same operation.

Real platforms still keep ranking separate, because they optimise for more than similarity: recency, engagement rate, promoted content, diversity, time of day, and signals from accounts a user follows. Those get modeled with a second, dedicated system, often a gradient-boosted tree or a neural network trained to predict engagement.

For a 500-video database built to test an idea, FAISS distance plus the 70-30 mix was enough. This does not prove a ranking layer is unnecessary in general. It shows that at this scale, the extra machinery would not have changed much, and adding it would have been effort spent on a problem the project did not actually have yet.

Limits

Small, synthetic database. 500 videos with text descriptions and pre-assigned categories, not real video, image or audio content. The embeddings describe text about a video, not the video itself.

No timestamps. The sliding window approximates time decay with a fixed window of 5 recent high-score videos rather than an actual decay curve, because interactions were never timestamped.

No ranking model. FAISS distance stands in for a proper ranking signal. It ignores recency, monetisation and diversity beyond the fixed 70-30 split.

Single user at a time. The demo tracks one user_embedding global variable. It shows the mechanism, not a multi-user production system.

What is not done yet

The next steps would be:

  1. Learn how to combine video embeddings into a user embedding, instead of a fixed weighted average, so the model can weigh interaction types and detect a genuine shift in taste versus a one-off click.
  2. Switch to IndexHNSWFlat, a graph-based FAISS index that supports adding new videos without rebuilding the whole index, so trending content can surface without a full reindex.
  3. Add real timestamps and replace the sliding window with exponential time decay.
  4. Add a diversity penalty (Maximal Marginal Relevance) so consecutive recommendations are pulled away from each other, not just from a 70-30 split.
  5. Embed actual video frames and audio, not text descriptions, using a vision model such as CLIP alongside the text embeddings.
  6. Add collaborative filtering, recommending based on what similar users like, not only the current user's own history.