How does a feed learn what you like?
I built a TikTok-style recommendation engine from scratch to answer a question I couldn't shake: does the app know a video is about cats, or does it just know I keep watching things like it? This page walks through the build, from a scoring rule you could run on paper to Google Gemini embeddings and a FAISS vector index.
Live at social-media-algo-code.onrender.com (a free Render instance, so the first load can take a moment to wake up). Code at github.com/ayushxpatne/social-media-algo-code.
- A scoring rule you could run by hand
- Teaching the system what a video is
- Searching millions of videos fast
- Keeping the feed fresh
- Why there is no separate ranking step
- Limits
- What is not done yet
A real run of the first version, ten interactions in
Only three topics exist: cats, dogs and space. A like scores +1, a dislike scores -1. The user saw 3 cat videos and liked 2, saw 5 dog videos and liked 4, saw 2 space videos and liked neither.
| Topic | Shown | Liked | Mean score |
|---|---|---|---|
| dogs | 5 | 4 | 0.60 |
| cats | 3 | 2 | 0.33 |
| space | 2 | 0 | -1.00 |
the dumb algorithmmean of plus-one, minus-one
preference_score = mean(scores per topic) dogs: 0.60 cats: 0.33 space: -1.00 next batch: mostly dogs, some cats, no space
This is the whole system at the start of the project: three topics, one score per interaction, and a mean. It ranks the topics correctly, but it does not know why a video is a dog video. Someone had to label it that way by hand. Sections below replace that label with an embedding, and the mean with a vector average.
A scoring rule you could run by hand
The first version has no machine learning in it at all. Every video already carries a hand-assigned topic. Each interaction adds a score to that topic, and a preference score is just the mean score per topic so far. The hero example above is a real run of it.
To turn that ranking into a feed, the batch of videos to show is split across topics in proportion to preference score:
With the scores above, a batch of 20 videos works out to about 13 dog videos and 7 cat videos, and zero space videos. That is exploitation: showing more of what a user already likes. Its opposite, exploration, means showing something new on purpose to learn whether the user likes it too. This first version does no exploration at all.
This rule works, but it does not scale. A real platform has thousands of categories, not three, and someone has to hand-label every video's topic before the mean can even be computed. Growing the interaction types helps a little:
interactions = {
# view_time under 2s counts as a skip, under 7s as a short watch, over 7s as a long watch
"view_time": {"skip": -1, "short": -0.5, "long": +1.5},
"like": +1,
"comment": +2,
"share": +3,
"save": +3,
}
More signals give a finer-grained score, but the core problem stays: the system still only knows a video's topic because a human typed it in. It has no way to look at a new video and place it anywhere.
Teaching the system what a video is
A machine learning model cannot read the word "cat". It only works with numbers. The fix is an embedding: a list of numbers that acts like a coordinate for a piece of content, placed so that similar things end up near each other.
Picture a toy space with three axes: how alive something looks, how cute or friendly it feels, and how mechanical it is. A dog might sit at (0.7, 0.9, 0.2), a cat close by at (0.8, 0.7, 0.3), and a car far off at (0.1, 0.2, 0.9). Nobody chose those numbers by hand. A neural network learns them by looking at huge amounts of text (or video, or audio) and gradually placing things that appear in similar contexts closer together.
So does the algorithm know a video is a cat video? Yes and no. It knows the video's position in a 3072-number space where cat videos cluster together. It has no label that reads "cats". In this project the inputs are still text, categories and descriptions converted to embeddings, not the raw video frames, but the principle is the one real platforms use on actual pixels and audio.
Building one number for a whole user
To recommend content, the system also needs a single point that represents what one user likes: a user embedding. It is built as a weighted average of the embeddings of videos that user engaged with, weighted by the score each video got.
Searching millions of videos fast
Once both videos and the user are points in the same space, finding recommendations means finding the video embeddings closest to the user embedding. Checking every video one by one, called a brute-force or flat search, works fine for 500 videos. A real platform has millions, and checking all of them for every user, every few seconds, does not scale.
FAISS (Facebook AI Similarity Search) is a library built for exactly this: it builds an index over the embeddings so the system can jump close to the answer instead of checking everything.
IndexFlatL2 brute force
index = faiss.IndexFlatL2(dimension)
for row in videos_db.itertuples():
index.add(row.embeddings.reshape(1, -1))
dist, ids = index.search(query, k=5)
Compares the query to every vector in the index and returns the k closest by Euclidean distance, the straight-line distance between two points. Exact, and correct, but a full scan every time. Fine at 500 rows, too slow at millions.
IndexIVFFlat clustered
quantizer = faiss.IndexFlatL2(dimensions) index = faiss.IndexIVFFlat(quantizer, dimensions, num_clusters) index.train(embeddings) index.add(embeddings) index.nprobe = 5 dist, ids = index.search(query, k=10)
First groups all embeddings into clusters with k-means, using a set number of centroids. A search only checks the nearest nprobe clusters instead of the whole index, trading a little accuracy for a lot of speed.
The trade-off shows up directly in one setting: nprobe, how many clusters a search is allowed to look inside. With the default of 1 cluster, asking for 100 similar videos out of a small cluster runs out and pads the result with -1, meaning "nothing left to return". Raising nprobe to 3 searches more clusters and returns a full, more varied list.
nprobe = 1 (default)
- What it shows
- A request for k=100 similar videos, searching only the single nearest cluster.
- What to look for
- Real ids for the first 82 results, then
-1padding for the rest. - What we saw
- The cluster held about 82 videos. Everything past that came back as
-1: the index had genuinely run out.
nprobe = 3
- What it shows
- The same request, now searching the 3 nearest clusters.
- What to look for
- Whether the
-1padding disappears and new ids show up further down the list. - What we saw
- A full 100 results with no padding, and more variety near the end of the list, at the cost of one extra cluster scan.
Keeping the feed fresh
A first working version of the FAISS-based feed had two problems. Videos arrived in clumps from the same topic, and after about 20 videos the feed would run out of similar items to show and stop. Two small changes fixed both.
A sliding window instead of a lifetime average
The user embedding was originally a weighted average of a user's entire watch history. That means someone who binged cat videos yesterday still gets a cat-heavy feed today even if they've since moved on. The fix is a sliding window: only the last 5 videos with a score of 2 or higher (meaning they were properly watched, not just skipped past) feed into the user embedding.
filtered_df = history_df[history_df["score"] >= 2].tail(5)
for row in filtered_df.itertuples():
embedding_matrix.append(row.embeddings)
weights.append(int(row.score) + 1)
weights = weights / sum(weights)
user_embedding = weighted_average(embedding_matrix, weights)
This is a stand-in for time decay, where older interactions matter less by a formula rather than a hard cutoff. The project does not record timestamps, so a fixed window of 5 recent, high-score videos is the simpler substitute: weight_i = score_i * e^(-lambda * time_elapsed_i) is the version that would replace it if timestamps were tracked.
The 70-30 mix
The running-out problem and the clumping problem share one fix. Instead of asking FAISS for 10 similar videos every batch, the system asks for only 7, then adds 3 videos picked at random from the whole database, and shuffles the combined 10 before sending them to the user. A video already shown gets swapped for a random one on the spot, so the feed never has fewer than 10 videos to show.
Try the mix yourself
Drag the slider to change how many of the next 10 videos come from the user's taste (FAISS neighbours of the user embedding, mint) versus picked at random for exploration (apricot). The project's own default is 7 similar to 3 random.
Why there is no separate ranking step
Every earlier version of this project sorted videos by preference score before serving them, a separate ranking step. Once FAISS entered the picture, that step quietly disappeared, because a similarity search already returns results ordered by distance: the closest video to the user embedding comes first. Retrieval and ranking collapsed into the same operation.
Real platforms still keep ranking separate, because they optimise for more than similarity: recency, engagement rate, promoted content, diversity, time of day, and signals from accounts a user follows. Those get modeled with a second, dedicated system, often a gradient-boosted tree or a neural network trained to predict engagement.
For a 500-video database built to test an idea, FAISS distance plus the 70-30 mix was enough. This does not prove a ranking layer is unnecessary in general. It shows that at this scale, the extra machinery would not have changed much, and adding it would have been effort spent on a problem the project did not actually have yet.
Limits
Small, synthetic database. 500 videos with text descriptions and pre-assigned categories, not real video, image or audio content. The embeddings describe text about a video, not the video itself.
No timestamps. The sliding window approximates time decay with a fixed window of 5 recent high-score videos rather than an actual decay curve, because interactions were never timestamped.
No ranking model. FAISS distance stands in for a proper ranking signal. It ignores recency, monetisation and diversity beyond the fixed 70-30 split.
Single user at a time. The demo tracks one user_embedding global variable. It shows the mechanism, not a multi-user production system.
What is not done yet
The next steps would be:
- Learn how to combine video embeddings into a user embedding, instead of a fixed weighted average, so the model can weigh interaction types and detect a genuine shift in taste versus a one-off click.
- Switch to
IndexHNSWFlat, a graph-based FAISS index that supports adding new videos without rebuilding the whole index, so trending content can surface without a full reindex. - Add real timestamps and replace the sliding window with exponential time decay.
- Add a diversity penalty (Maximal Marginal Relevance) so consecutive recommendations are pulled away from each other, not just from a 70-30 split.
- Embed actual video frames and audio, not text descriptions, using a vision model such as CLIP alongside the text embeddings.
- Add collaborative filtering, recommending based on what similar users like, not only the current user's own history.