# Semantic Search Over Transcripts in Python: What Embeddings Fix and What They Don't

Short answer: keyword search finds the words you typed, and vector search finds paragraphs that sit close to them in meaning. I built both in about 120 lines of Python, ran them on a made-up talk transcript, and counted the results. Vector search found more of the right paragraphs for one query and none for the other. This post shows the code, the real output, and why.

One note before you start: the vector search in this post is LSA, a small stand-in that learns only from the 28 paragraphs. A pretrained model can behave differently, and I did not run one for the results below.

I build Libraryminds, a tool that makes video searchable, so this problem is close to my work. The code here does not use it. It runs on your machine with scikit-learn and nothing to download.

## What I tested

I wrote a 50 minute transcript of an imaginary talk about running a small remote team. It has 28 paragraphs, each with a timestamp. It is made up, so this is a demo and not a benchmark.

I picked two queries, and I wrote down which paragraphs really answer them before I ran any search:

1.  "team morale" has 5 right answers. Only one of them uses the word morale. The others talk about burnout, tiredness, overtime, energy and motivation.
    
2.  "reduce churn" has 3 right answers. They are about people who cancel in the first month. The word churn is never said. Another paragraph uses the word reduce, about meetings, as a trap.
    

I set the trap on purpose. Real speakers say "cancel" and "leave" far more often than "churn".

## Step 1: One result is one paragraph with a timestamp

A search result has to be something a person can act on. So the transcript is a list of (timestamp, paragraph) pairs. You search the paragraphs, and you show the timestamp so the reader can jump to the video and check.

## Step 2: Keyword search with BM25

```python
# 1. Keyword search: BM25 over exact words. No stemming, no synonyms.
bm25 = BM25Okapi([tokens(t) for t in TEXTS])


def keyword_search(query, top=5):
    scores = bm25.get_scores(tokens(query))
    order = np.argsort(-scores)
    return [(int(i), float(scores[i])) for i in order[:top] if scores[i] > 0]
```

BM25 scores a paragraph higher when it contains your words, and more so when the word is rare in the transcript. It matches exact words only. There is no stemming, so "cancel" does not match "cancellations", and there are no synonyms.

## Step 3: Vector search

```python
# 2. Vector search: TF-IDF, then SVD to 8 dimensions (latent semantic analysis).
tfidf = TfidfVectorizer(stop_words="english", sublinear_tf=True)
svd = TruncatedSVD(n_components=8, random_state=0)
DOC_VECS = normalize(svd.fit_transform(tfidf.fit_transform(TEXTS)))


def embed(query):
    return normalize(svd.transform(tfidf.transform([query])))[0]


def vector_search(query, top=5):
    q = embed(query)
    if not q.any():
        return []  # none of the query words exist in the corpus
    sims = DOC_VECS @ q
    order = np.argsort(-sims)
    return [(int(i), float(sims[i])) for i in order[:top] if sims[i] > 0]
```

An embedding turns a paragraph into a list of numbers, so that paragraphs with similar meaning get similar numbers. Search then means turning the query into numbers and finding the closest paragraphs by cosine similarity.

Here the embedding is TF-IDF squeezed into 8 dimensions with SVD. This is called latent semantic analysis. It is old and small, and it needs no download, which is why I used it. It has one big limit: it learns only from these 28 paragraphs. If two words never appear near each other in the transcript, it cannot know they are related.

The function embed() is the part you can swap later.

## Step 4: Join the two lists

```python
# 3. Hybrid: Reciprocal Rank Fusion. Each list gives 1 / (60 + rank) to a paragraph.
def hybrid_search(query, top=5, k=60):
    fused = {}
    for results in (keyword_search(query, 20), vector_search(query, 20)):
        for rank, (i, _) in enumerate(results, start=1):
            fused[i] = fused.get(i, 0.0) + 1.0 / (k + rank)
    ranked = sorted(fused.items(), key=lambda x: -x[1])
    return ranked[:top]
```

This is reciprocal rank fusion. Each list gives a paragraph 1 divided by (60 plus its rank), and the scores are added. Keyword scores and vector scores are on different scales, so you cannot add them directly, but ranks you can. The 60 is the value most people use.

## The real output

The output below is from a real run, with Python 3.13, scikit-learn 1.9.1 and numpy 2.5.3. A star marks a paragraph that really answers the query.

```text
Query: "team morale"   (* = really answers the query)
  keyword  found 2 of 5 relevant in top 5
   * 00:07:40  4.455  Now the part people ask about most, team morale. Honestly ...
     00:00:25  1.897  Welcome everyone. I run a small engineering team of nine p...
   * 00:11:00  1.780  We fixed it with one boring change: no messages after seve...
  vector   found 3 of 5 relevant in top 5
   * 00:07:40  0.929  Now the part people ask about most, team morale. Honestly ...
   * 00:14:05  0.724  Overtime is a loan with a high interest rate. You get the ...
     00:00:25  0.624  Welcome everyone. I run a small engineering team of nine p...
   * 00:11:00  0.561  We fixed it with one boring change: no messages after seve...
     00:36:20  0.497  We rewrote the first week. Three short emails, one sample ...
  hybrid   found 3 of 5 relevant in top 5
   * 00:07:40  0.033  Now the part people ask about most, team morale. Honestly ...
     00:00:25  0.032  Welcome everyone. I run a small engineering team of nine p...
   * 00:11:00  0.031  We fixed it with one boring change: no messages after seve...
   * 00:14:05  0.016  Overtime is a loan with a high interest rate. You get the ...
     00:36:20  0.015  We rewrote the first week. Three short emails, one sample ...

Query: "reduce churn"   (* = really answers the query)
  keyword  found 0 of 3 relevant in top 5
     00:29:50  2.872  Meetings were eating our days. We set a limit of two per p...
  vector   found 0 of 3 relevant in top 5
     00:29:50  0.985  Meetings were eating our days. We set a limit of two per p...
     00:02:10  0.737  We started with one rule: write things down. If a decision...
     00:23:10  0.688  Incidents taught us the most. Our rule is that the first m...
     00:16:20  0.579  Onboarding used to take a month. Now a new person ships a ...
     00:26:40  0.499  On call is shared by four people, one week each. Whoever i...
  hybrid   found 0 of 3 relevant in top 5
     00:29:50  0.033  Meetings were eating our days. We set a limit of two per p...
     00:02:10  0.016  We started with one rule: write things down. If a decision...
     00:23:10  0.016  Incidents taught us the most. Our rule is that the first m...
     00:16:20  0.016  Onboarding used to take a month. Now a new person ships a ...
     00:26:40  0.015  On call is shared by four people, one week each. Whoever i...
```

## What the output says

1.  For "team morale", keyword search found 2 of 5 right paragraphs and vector search found 3 of 5. The extra one is 00:14:05, the overtime paragraph. It has none of the query words, and vector search still found it. That is the gain.
    
2.  Hybrid also found 3 of 5, the same as vector. Joining the lists did not beat vector search here. With 28 paragraphs there is little room for it to help, so I would not claim hybrid is better from this test.
    
3.  Vector search also returned paragraphs that are not right, like 00:00:25 (the welcome, which has the word team) and 00:36:20 (the onboarding emails). It does not say "I am not sure". It just ranks them.
    
4.  For "reduce churn", all three methods found 0 of 3. The word churn is not in the transcript, so this small model has no way to link it to "cancel". And the word reduce pulled in the meetings paragraph, which all three methods ranked first. It is confident and wrong.
    

## How much do the settings matter?

I chose 8 dimensions before the first run. After seeing the results I added one more test, to see how much that number matters:

```python
# How much does the number of dimensions matter? (I added this after the first run.)
print("\nVector search, top 5, with a different number of dimensions:")
X = tfidf.transform(TEXTS)
for n in (2, 4, 6, 8, 12, 16):
    s = TruncatedSVD(n_components=n, random_state=0)
    docs = normalize(s.fit_transform(X))
    row = []
    for query in RELEVANT:
        q = normalize(s.transform(tfidf.transform([query])))[0]
        sims = docs @ q if q.any() else np.zeros(len(TEXTS))
        top = [i for i in np.argsort(-sims)[:5] if sims[i] > 0]
        row.append(len({TIMES[i] for i in top} & RELEVANT[query]))
    print(f"  {n:2} dimensions: morale {row[0]} of 5, churn {row[1]} of 3")
```

```text
Vector search, top 5, with a different number of dimensions:
   2 dimensions: morale 2 of 5, churn 0 of 3
   4 dimensions: morale 3 of 5, churn 1 of 3
   6 dimensions: morale 3 of 5, churn 0 of 3
   8 dimensions: morale 3 of 5, churn 0 of 3
  12 dimensions: morale 3 of 5, churn 0 of 3
  16 dimensions: morale 3 of 5, churn 0 of 3
```

For "team morale" the result is steady at 3 of 5 from 4 dimensions up, and drops to 2 of 5 at 2 dimensions. For "reduce churn" it is 0 almost everywhere. The single 1 at 4 dimensions looks like luck to me, not skill, so I would not build on it.

## Swap in a pretrained model

This part was not part of the run above. A pretrained model has read far more text than 28 paragraphs, so it does not need your transcript to put "churn" near "cancel". The two functions in this post only need DOC\_VECS and embed(), so the swap is small:

```python
# pip install sentence-transformers
from sentence_transformers import SentenceTransformer

model = SentenceTransformer("all-MiniLM-L6-v2")
DOC_VECS = model.encode(TEXTS, normalize_embeddings=True)

def embed(query):
    return model.encode([query], normalize_embeddings=True)[0]
```

Put it after the TEXTS list and run the script again. Then compare the new numbers with the table above. Whether it fixes the "reduce churn" query on this transcript is now something you can measure instead of assume.

## How a real product describes it

On its feature page, Libraryminds describes its semantic search like this: vector embeddings made from the transcript text, results ranked by similarity at paragraph level, keyword and semantic modes side by side, and results in under a second because the embeddings are computed in advance. It is on the Plus plan and above. Those are claims from the [semantic search feature page](https://libraryminds.com/features/ai-semantic-search). This demo does not test them.

## What to take from this

1.  Return a paragraph with a timestamp, not a whole document.
    
2.  Keep keyword search for names, numbers and exact quotes. Vectors can blur them.
    
3.  Before you trust vector search, write 5 to 10 test queries and the right answers, then count.
    
4.  Expect confident wrong results. Show the timestamp so a person can check.
    
5.  Treat the embedding as a part you can replace, and measure again when you do.
    

## The full script

Install the packages with `pip install scikit-learn rank-bm25 numpy`, save this as search\_demo.py, and run `python search_demo.py`.

```python
import re
import numpy as np
from rank_bm25 import BM25Okapi
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.decomposition import TruncatedSVD
from sklearn.preprocessing import normalize

# A made-up transcript of a talk. Each item is (timestamp, paragraph).
TRANSCRIPT = [
    ("00:00:25", "Welcome everyone. I run a small engineering team of nine people spread over four time zones, and today I want to share what we learned in our second year."),
    ("00:02:10", "We started with one rule: write things down. If a decision was made in a call, it goes into a shared document the same day, with the name of the person who made it."),
    ("00:04:05", "Our standup is now written, not spoken. Each person posts three lines in the morning, and we only meet if somebody is blocked."),
    ("00:05:50", "Hiring was our biggest bet. We stopped asking for puzzle questions and started giving a paid two day task that looks like real work."),
    ("00:07:40", "Now the part people ask about most, team morale. Honestly it fell in the spring, and the cause was burnout, not pay: people were tired and they stopped caring about small details."),
    ("00:09:15", "When I looked at the calendar I saw the pattern. Three people had worked late every night for six weeks, and nobody had said a word because the launch felt important."),
    ("00:11:00", "We fixed it with one boring change: no messages after seven in the evening, and a real week of rest after every release. The energy in the team came back within a month."),
    ("00:12:30", "Motivation is not a speech. It comes from finishing things, so we cut the roadmap in half and shipped smaller pieces more often."),
    ("00:14:05", "Overtime is a loan with a high interest rate. You get the feature this week and you pay for it with mistakes and resignations later."),
    ("00:16:20", "Onboarding used to take a month. Now a new person ships a tiny change on day two, because the setup script and the checklist are always up to date."),
    ("00:18:00", "Documentation rots unless somebody owns it. Every page has an owner and a review date, and the date shows at the top of the page."),
    ("00:19:45", "On pricing, we ran only two plans for the first year. Fewer choices meant fewer questions, and our support inbox got quieter."),
    ("00:21:30", "We tried usage based pricing for a quarter and dropped it. Customers could not guess their own bill, and that fear cost us trials."),
    ("00:23:10", "Incidents taught us the most. Our rule is that the first message goes out in ten minutes, even if all it says is that we are looking into it."),
    ("00:25:00", "After every incident we write a short review with no names in it. The goal is to find the missing alarm or the unclear step, not the guilty person."),
    ("00:26:40", "On call is shared by four people, one week each. Whoever is on call can skip meetings, and that rule alone saved a lot of sleep."),
    ("00:28:15", "Deployments happen every day before lunch. A small change is easy to roll back, and nobody has to stay late to watch it."),
    ("00:29:50", "Meetings were eating our days. We set a limit of two per person per day, and we reduce the length of the rest to twenty five minutes."),
    ("00:31:10", "Now customers. A large share of people cancel in the first month, and for a long time we did not know why they left."),
    ("00:33:45", "So we called forty of them. Most said the same thing: they signed up with a big plan in mind and never got past the first import."),
    ("00:36:20", "We rewrote the first week. Three short emails, one sample project already loaded, and a person to reply within an hour. Cancellations in the first month fell noticeably."),
    ("00:38:00", "Customer interviews changed our roadmap. We ask three questions: what were you trying to do, what got in the way, and what did you do instead."),
    ("00:40:30", "Security is mostly habit. Password manager for everybody, hardware keys for admins, and a yearly check of who still has access to what."),
    ("00:42:10", "We keep a small monthly budget for experiments. Each experiment gets a written question, a date to stop, and one person who decides."),
    ("00:44:00", "Retros are short. Three columns on a page: keep, stop, try, and every try has a name next to it by the end of the call."),
    ("00:46:30", "Remote work is not the hard part. The hard part is trust, and trust comes from small promises that are kept, again and again."),
    ("00:48:20", "If I could give one piece of advice, it would be to protect your people's time and attention before you protect your roadmap."),
    ("00:50:00", "Thank you for listening. The notes and the checklists from this talk are in the shared folder, and I will answer questions now."),
]

# Which paragraphs really answer each query. I decided this when I wrote the
# transcript, before I ran any search.
RELEVANT = {
    "team morale": {"00:07:40", "00:09:15", "00:11:00", "00:12:30", "00:14:05"},
    "reduce churn": {"00:31:10", "00:33:45", "00:36:20"},
}

TIMES = [t for t, _ in TRANSCRIPT]
TEXTS = [p for _, p in TRANSCRIPT]


def tokens(text):
    return re.findall(r"[a-z']+", text.lower())


# 1. Keyword search: BM25 over exact words. No stemming, no synonyms.
bm25 = BM25Okapi([tokens(t) for t in TEXTS])


def keyword_search(query, top=5):
    scores = bm25.get_scores(tokens(query))
    order = np.argsort(-scores)
    return [(int(i), float(scores[i])) for i in order[:top] if scores[i] > 0]


# 2. Vector search: TF-IDF, then SVD to 8 dimensions (latent semantic analysis).
tfidf = TfidfVectorizer(stop_words="english", sublinear_tf=True)
svd = TruncatedSVD(n_components=8, random_state=0)
DOC_VECS = normalize(svd.fit_transform(tfidf.fit_transform(TEXTS)))


def embed(query):
    return normalize(svd.transform(tfidf.transform([query])))[0]


def vector_search(query, top=5):
    q = embed(query)
    if not q.any():
        return []  # none of the query words exist in the corpus
    sims = DOC_VECS @ q
    order = np.argsort(-sims)
    return [(int(i), float(sims[i])) for i in order[:top] if sims[i] > 0]


# 3. Hybrid: Reciprocal Rank Fusion. Each list gives 1 / (60 + rank) to a paragraph.
def hybrid_search(query, top=5, k=60):
    fused = {}
    for results in (keyword_search(query, 20), vector_search(query, 20)):
        for rank, (i, _) in enumerate(results, start=1):
            fused[i] = fused.get(i, 0.0) + 1.0 / (k + rank)
    ranked = sorted(fused.items(), key=lambda x: -x[1])
    return ranked[:top]


def show(name, results, query, cut=5):
    hits = [TIMES[i] for i, _ in results[:cut]]
    found = len(set(hits) & RELEVANT[query])
    print(f"  {name:8} found {found} of {len(RELEVANT[query])} relevant in top {cut}")
    for i, s in results[:cut]:
        mark = "*" if TIMES[i] in RELEVANT[query] else " "
        print(f"   {mark} {TIMES[i]}  {s:.3f}  {TEXTS[i][:58]}...")
    if not results:
        print("     (no results)")


for query in RELEVANT:
    print(f'\nQuery: "{query}"   (* = really answers the query)')
    show("keyword", keyword_search(query), query)
    show("vector", vector_search(query), query)
    show("hybrid", hybrid_search(query), query)


# How much does the number of dimensions matter? (I added this after the first run.)
print("\nVector search, top 5, with a different number of dimensions:")
X = tfidf.transform(TEXTS)
for n in (2, 4, 6, 8, 12, 16):
    s = TruncatedSVD(n_components=n, random_state=0)
    docs = normalize(s.fit_transform(X))
    row = []
    for query in RELEVANT:
        q = normalize(s.transform(tfidf.transform([query])))[0]
        sims = docs @ q if q.any() else np.zeros(len(TEXTS))
        top = [i for i in np.argsort(-sims)[:5] if sims[i] > 0]
        row.append(len({TIMES[i] for i in top} & RELEVANT[query]))
    print(f"  {n:2} dimensions: morale {row[0]} of 5, churn {row[1]} of 3")
```

If you find a query where this breaks differently on your own transcripts, tell me in the comments.

## Update: what is vocabulary and what is the model?

A reader, Ahmet Özel, pointed out that the "reduce churn" miss is about vocabulary coverage in this LSA version, and not a general failure of semantic embeddings. He is right. In the LSA version, "churn" has no feature at all, so the query effectively becomes "reduce" before the projection.

He suggested printing which query words the model used, and comparing "reduce churn" with "reduce cancellations" and with a paraphrase that shares no words, while keeping the right answers fixed. Here is that test for keyword search and LSA. The right answers are the same 3 paragraphs for every query.

```text
* = one of the 3 right paragraphs. The right answers are the same for every query.

Query: "reduce churn"
  used by LSA: ['reduce']   stop words dropped: []   not in the transcript: ['churn']
  keyword      found 0 of 3 in top 5:   00:29:50
  LSA vector   found 0 of 3 in top 5:   00:29:50   00:02:10   00:23:10   00:16:20   00:26:40
  pretrained   not run (ModuleNotFoundError)

Query: "reduce cancellations"
  used by LSA: ['reduce', 'cancellations']   stop words dropped: []   not in the transcript: []
  keyword      found 1 of 3 in top 5:   00:29:50  *00:36:20
  LSA vector   found 1 of 3 in top 5:  *00:36:20   00:29:50   00:11:00   00:16:20   00:25:00
  pretrained   not run (ModuleNotFoundError)

Query: "subscriber attrition"
  used by LSA: []   stop words dropped: []   not in the transcript: ['subscriber', 'attrition']
  keyword      found 0 of 3 in top 5:  (no results)
  LSA vector   found 0 of 3 in top 5:  (no results)
  pretrained   not run (ModuleNotFoundError)
```

What it shows:

1.  "reduce churn": only "reduce" is used, so both searches find 0 of 3, and the meetings paragraph comes first.
    
2.  "reduce cancellations": both words are used. Both searches find 1 of 3. LSA puts the right paragraph first, and keyword search puts it second. The other two right paragraphs say "cancel", which is a different word, and neither search links it to "cancellations".
    
3.  "subscriber attrition": no word is in the transcript, so both searches return nothing.
    

In this small test, vocabulary changes the result the most. The learned projection changed the order once and the count never, and with 3 queries I would not read more into that.

I have not run the pretrained half. The model download was blocked where I ran this, so the pretrained lines say "not run". If you run the script below with `pip install sentence-transformers`, the third line of each query will fill in. I will add those numbers here when I have them.

```python
import re
import numpy as np
from rank_bm25 import BM25Okapi
from sklearn.feature_extraction.text import TfidfVectorizer, ENGLISH_STOP_WORDS
from sklearn.decomposition import TruncatedSVD
from sklearn.preprocessing import normalize

# A made-up transcript of a talk. Each item is (timestamp, paragraph).
TRANSCRIPT = [
    ("00:00:25", "Welcome everyone. I run a small engineering team of nine people spread over four time zones, and today I want to share what we learned in our second year."),
    ("00:02:10", "We started with one rule: write things down. If a decision was made in a call, it goes into a shared document the same day, with the name of the person who made it."),
    ("00:04:05", "Our standup is now written, not spoken. Each person posts three lines in the morning, and we only meet if somebody is blocked."),
    ("00:05:50", "Hiring was our biggest bet. We stopped asking for puzzle questions and started giving a paid two day task that looks like real work."),
    ("00:07:40", "Now the part people ask about most, team morale. Honestly it fell in the spring, and the cause was burnout, not pay: people were tired and they stopped caring about small details."),
    ("00:09:15", "When I looked at the calendar I saw the pattern. Three people had worked late every night for six weeks, and nobody had said a word because the launch felt important."),
    ("00:11:00", "We fixed it with one boring change: no messages after seven in the evening, and a real week of rest after every release. The energy in the team came back within a month."),
    ("00:12:30", "Motivation is not a speech. It comes from finishing things, so we cut the roadmap in half and shipped smaller pieces more often."),
    ("00:14:05", "Overtime is a loan with a high interest rate. You get the feature this week and you pay for it with mistakes and resignations later."),
    ("00:16:20", "Onboarding used to take a month. Now a new person ships a tiny change on day two, because the setup script and the checklist are always up to date."),
    ("00:18:00", "Documentation rots unless somebody owns it. Every page has an owner and a review date, and the date shows at the top of the page."),
    ("00:19:45", "On pricing, we ran only two plans for the first year. Fewer choices meant fewer questions, and our support inbox got quieter."),
    ("00:21:30", "We tried usage based pricing for a quarter and dropped it. Customers could not guess their own bill, and that fear cost us trials."),
    ("00:23:10", "Incidents taught us the most. Our rule is that the first message goes out in ten minutes, even if all it says is that we are looking into it."),
    ("00:25:00", "After every incident we write a short review with no names in it. The goal is to find the missing alarm or the unclear step, not the guilty person."),
    ("00:26:40", "On call is shared by four people, one week each. Whoever is on call can skip meetings, and that rule alone saved a lot of sleep."),
    ("00:28:15", "Deployments happen every day before lunch. A small change is easy to roll back, and nobody has to stay late to watch it."),
    ("00:29:50", "Meetings were eating our days. We set a limit of two per person per day, and we reduce the length of the rest to twenty five minutes."),
    ("00:31:10", "Now customers. A large share of people cancel in the first month, and for a long time we did not know why they left."),
    ("00:33:45", "So we called forty of them. Most said the same thing: they signed up with a big plan in mind and never got past the first import."),
    ("00:36:20", "We rewrote the first week. Three short emails, one sample project already loaded, and a person to reply within an hour. Cancellations in the first month fell noticeably."),
    ("00:38:00", "Customer interviews changed our roadmap. We ask three questions: what were you trying to do, what got in the way, and what did you do instead."),
    ("00:40:30", "Security is mostly habit. Password manager for everybody, hardware keys for admins, and a yearly check of who still has access to what."),
    ("00:42:10", "We keep a small monthly budget for experiments. Each experiment gets a written question, a date to stop, and one person who decides."),
    ("00:44:00", "Retros are short. Three columns on a page: keep, stop, try, and every try has a name next to it by the end of the call."),
    ("00:46:30", "Remote work is not the hard part. The hard part is trust, and trust comes from small promises that are kept, again and again."),
    ("00:48:20", "If I could give one piece of advice, it would be to protect your people's time and attention before you protect your roadmap."),
    ("00:50:00", "Thank you for listening. The notes and the checklists from this talk are in the shared folder, and I will answer questions now."),
]


# The right answers are fixed for every query variant below. They are the three
# paragraphs about people who cancel in the first month.
RELEVANT = {"00:31:10", "00:33:45", "00:36:20"}

QUERIES = [
    "reduce churn",              # "churn" is not in the transcript
    "reduce cancellations",      # both words are in the transcript
    "subscriber attrition",      # no word in the transcript at all
]

TIMES = [t for t, _ in TRANSCRIPT]
TEXTS = [p for _, p in TRANSCRIPT]


def tokens(text):
    return re.findall(r"[a-z']+", text.lower())


# 1. Keyword search: BM25 over exact words.
bm25 = BM25Okapi([tokens(t) for t in TEXTS])


def keyword_ranking(query):
    scores = bm25.get_scores(tokens(query))
    return [int(i) for i in np.argsort(-scores) if scores[i] > 0]


# 2. Vector search with LSA: TF-IDF, then SVD to 8 dimensions.
tfidf = TfidfVectorizer(stop_words="english", sublinear_tf=True)
svd = TruncatedSVD(n_components=8, random_state=0)
DOC_VECS = normalize(svd.fit_transform(tfidf.fit_transform(TEXTS)))


def lsa_ranking(query):
    q = normalize(svd.transform(tfidf.transform([query])))[0]
    if not q.any():
        return []
    sims = DOC_VECS @ q
    return [int(i) for i in np.argsort(-sims) if sims[i] > 0]


def query_terms(query):
    """Split the query words into: used by the LSA model, stop words, and unknown words."""
    used, stop, unknown = [], [], []
    for word in re.findall(r"\w\w+", query.lower()):
        if word in ENGLISH_STOP_WORDS:
            stop.append(word)
        elif word in tfidf.vocabulary_:
            used.append(word)
        else:
            unknown.append(word)
    return used, stop, unknown


# 3. Optional: a pretrained model. This part needs `pip install sentence-transformers`
#    and a first-time model download, so it is skipped when that fails.
try:
    from sentence_transformers import SentenceTransformer
    model = SentenceTransformer("all-MiniLM-L6-v2")
    PRE_DOCS = model.encode(TEXTS, normalize_embeddings=True)

    def pretrained_ranking(query):
        q = model.encode([query], normalize_embeddings=True)[0]
        sims = PRE_DOCS @ q
        return [int(i) for i in np.argsort(-sims)]

    PRETRAINED_NOTE = None
except Exception as error:
    pretrained_ranking = None
    PRETRAINED_NOTE = f"not run ({type(error).__name__})"


def line(name, ranking, top=5):
    if ranking is None:
        return f"  {name:12} {PRETRAINED_NOTE}"
    shown = ranking[:top]
    found = len({TIMES[i] for i in shown} & RELEVANT)
    marks = "  ".join(("*" if TIMES[i] in RELEVANT else " ") + TIMES[i] for i in shown) or "(no results)"
    return f"  {name:12} found {found} of {len(RELEVANT)} in top {top}:  {marks}"


print("* = one of the 3 right paragraphs. The right answers are the same for every query.\n")
for query in QUERIES:
    used, stop, unknown = query_terms(query)
    print(f'Query: "{query}"')
    print(f"  used by LSA: {used}   stop words dropped: {stop}   not in the transcript: {unknown}")
    print(line("keyword", keyword_ranking(query)))
    print(line("LSA vector", lsa_ranking(query)))
    print(line("pretrained", pretrained_ranking(query) if pretrained_ranking else None))
    print()
```
