Kurs
Manchmal reicht eine einfache Volltextsuche oder nur die Vektorsuche allein nicht aus, um eine Datenbank sinnvoll abzufragen und die gewünschten Ergebnisse zu erhalten. Die Kombination aus beidem ist ideal, wenn Entwickler mit großen Mengen multimodaler, unstrukturierter Daten arbeiten, die von beiden Sucharten profitieren. Das nennt sich Hybridsuche und bietet eine starke Lösung für eine knifflige Aufgabe.
Was genau ist Hybridsuche?
Um Hybridsuche richtig zu verstehen, müssen wir zuerst klären, was Volltextsuche ist.
Die Volltextsuche gleicht wörtliche Begriffe aus deiner Anfrage mit den Inhalten deiner Dokumente ab. Diese klassische Suche ist vielen Entwicklern bestens vertraut.
Suchst du zum Beispiel nach „gemütliches Café mit Außensitzplätzen“, prüft die Suchmaschine genau diese Wörter in der Datenbank. Kurz gesagt: Volltextsuche ist sehr präzise und effizient, funktioniert aber schlecht, wenn du Synonyme, Paraphrasen verwendest oder Tippfehler in der Anfrage hast.
Die Vektorsuche hingegen wandelt alle Daten in Zahlen bzw. Embeddings um. Statt exakte Wörter zu matchen, vergleicht die Vektorsuche die semantische Bedeutung deiner Anfrage mit den Dokumenten in der Datenbank.
Die Suche nach „gemütliches Café mit Außensitzplätzen“ kann so „Gebäck und Kaffee draußen“ ergeben – auch wenn die Wörter nicht exakt übereinstimmen. Vektorsuche ist nicht nur semantisch, sondern auch sehr flexibel, kann jedoch je nach Anfrage manchmal zu breit streuen.
Und wo kommt nun die Hybridsuche ins Spiel? Sie kombiniert Volltext- und Vektorsuche. So können Entwickler die semantische Intelligenz von Vektoren nutzen und gleichzeitig die präzisen Filtermöglichkeiten der Volltextsuche beibehalten. Das Beste aus beiden Welten – besonders hilfreich bei großen, unstrukturierten Datensätzen.
Warum Hybridsuche wichtig ist
Hybridsuche zahlt sich in vielen realen Anwendungsfällen aus, unter anderem im E-Commerce, Gesundheitswesen oder Recruiting.
Im E-Commerce stell dir vor, du suchst auf deiner Lieblingswebsite nach „bequemer Bürostuhl unter 100 $“. Die Vektorsuche zeigt dir semantisch ähnliche Produkte (z. B. ergonomische Bürostühle), während die Volltextsuche auf „Stuhl“ fokussiert und gleichzeitig die Preisgrenze durchsetzt.
Im Recruiting: Sucht ein Recruiter in Lebensläufen nach „Engineer mit NLP-Erfahrung“, erfasst die Hybridsuche sowohl „Natural Language Processing“ als auch das exakte Stichwort „Engineer“.
So erhältst du relevantere und verlässlichere Ergebnisse als mit jeder einzelnen Methode für sich.
Hybridsuche in MongoDB
Schauen wir uns an, wie Hybridsuche in MongoDB Atlas funktioniert. Für dieses Tutorial brauchst du ein paar Voraussetzungen:
In diesem Tutorial arbeiten wir mit der Collection „embedded movies“ in der Datenbank sample_mflix und orientieren uns am $rankFusion Hybridsuche-Tutorial – mit ein paar Anpassungen.
Der neue $rankFusion-Operator ist in Clustern ab Version 8.1+ verfügbar und vereinfacht Hybridsuche in MongoDB enorm, befindet sich aber noch in der Public Preview. Free-Tier-Cluster laufen automatisch auf Version 8.0. Wir folgen daher dem Tutorial, passen ein paar Dinge an und bringen es sauber zu Ende.
Cluster bereitstellen
Stelle beim Einrichten deines Clusters sicher, dass du den Datensatz sample_mflix lädst. Mit diesem Datensatz arbeiten wir im gesamten Tutorial. Unser Fokus liegt auf der Collection sample_mflix.embedded_movies.

Prüfe kurz, ob mongosh installiert ist.
mongosh --version
Für dieses Tutorial nutze ich Version 2.4.2.
Verbinde dich jetzt mit deinem MongoDB Atlas Cluster:
mongosh "mongodb+srv://<yourclusterhere>.mongodb.net/" --apiVersion 1 --username <yourusername>
Du wirst nach dem Passwort gefragt und siehst nach der Verbindung etwa Folgendes:

Jetzt können wir unsere Datenbank auswählen. Führe aus:
use sample_mflix
Die Ausgabe lautet:
switched to db sample_mflix
Indizes erstellen
Jetzt erstellen wir die beiden Indizes, die wir benötigen: den Vektorindex und den Volltextindex. Beide legen wir auf der Collection embedded_movies an.
Vektorindex:
db.embedded_movies.createSearchIndex(
"hybrid-vector-search",
"vectorSearch",
{
fields: [
{ type: "vector", path: "plot_embedding_voyage_3_large", numDimensions: 2048, similarity: "dotProduct" }
]
}
)
Volltextindex:
db.embedded_movies.createSearchIndex(
"hybrid-full-text-search",
"search",
{ mappings: { dynamic: true } }
)
Prüfe im MongoDB Atlas Cluster, ob die Indizes bereit sind:

Datenbank abfragen
Wir wollen nun die Daten in sample_mflix.embedded_movies für „star wars“ im Feld plot_embedding_voyage_3_large abfragen.
Da wir keine APIs verwenden müssen, können wir alle benötigten Embeddings in einer separaten Datei namens query_embeddings.js speichern.

Lade die Embeddings, um sie in der Abfrage zu verwenden. Führe Folgendes aus:
load('/Users/<PATH NAME>/query_embeddings.js')
const queryVec = STAR_WARS_EMBEDDING
Um zu prüfen, ob die Embeddings geladen wurden, führe aus:
STAR_WARS_EMBEDDING.length

Eine Ausgabe von 2048 ist korrekt.
Aggregation-Pipeline
Jetzt erstellen wir eine Aggregation-Pipeline für die Hybridsuche. Beachte, dass es ein paar Einschränkungen gibt, die unseren Ansatz bestimmen.
Wie oben erwähnt, ist $rankFusion in MongoDB Atlas zwar verfügbar, derzeit aber in Public Preview und erfordert ein Cluster 8.0 oder höher. Mit einem Free-Tier-Cluster kannst du $rankFusion aktuell nicht nutzen. Dadurch fusioniert Atlas die Text- und Vektorrankings nicht nativ. Wir müssen also selbst fusionieren – sprich, eine gewichtete Mischung beider Ranglisten bilden.
Die Positionen der Stufen in Aggregation-Pipelines sind ebenfalls streng:
$searchmuss die erste Stufe der Pipeline sein.$vectorSearchmuss den Beginn einer Pipeline bilden und darf nicht in$facetstehen.
Der Workaround: $vectorSearch kann die erste Stufe einer Sub-Pipeline innerhalb von $unionWith sein, da diese Sub-Pipeline als eigene Pipeline zählt.
Deshalb starten wir mit $search, damit die Volltextergebnisse ihren Score behalten, und nutzen dann $unionWith als zweite Mini-Pipeline, die mit $vectorSearch beginnt – so bleibt auch dessen Score erhalten.
Anschließend gruppieren wir nach _id, um Duplikate zu vermeiden und jeweils das beste Ergebnis aus $search und $vectorSearch zu nutzen.
Dann bestimmen wir die Hybridsuche: den besten $search-Score plus den besten $vectorSearch-Score, multipliziert mit einem frei wählbaren Gewicht (hier 0,35), und sortieren danach.
Das Gewicht legt fest, wie stark die Vektorsuche gegenüber der Volltextsuche die Endergebnisse beeinflusst.
Unsere Hybrid-Formel lautet: hybrid_score = text_score + (vector_score x weight). Ein höheres Gewicht (z. B. 0,7–1,0) lässt die Vektorsuche dominieren und priorisiert semantische Ähnlichkeit.
Ein niedrigeres Gewicht (0,1–0,3) bedeutet dagegen, dass die Volltextsuche dominiert und exakte Keyword-Treffer bevorzugt.
Kopiere diese Aggregation-Pipeline in dein Terminal:
const WEIGHT = 0.35;
db.embedded_movies.aggregate([
// Stage 1: full-text search
{ $search: { index: "hybrid-full-text-search", text: { query: "star wars", path: ["title","plot"] } } },
{ $set: { t: { $meta: "searchScore" } } },
{ $project: { _id: 1, title: 1, plot: 1, t: 1 } },
// Stage 2: Union with vector search
{ $unionWith: {
coll: "embedded_movies",
pipeline: [
// Vector search subpipeline
{ $vectorSearch: {
index: "hybrid-vector-search",
path: "plot_embedding_voyage_3_large",
queryVector: queryVec,
numCandidates: 200,
limit: 150
}},
{ $project: { _id: 1, title: 1, plot: 1, v: { $meta: "vectorSearchScore" } } }
]
}},
// Stage 3: Combine and rank results
{ $group: { _id: "$_id", title: { $first: "$title" }, plot: { $first: "$plot" }, t: { $max: "$t" }, v: { $max: "$v" } } },
{ $set: { h: { $add: [ { $ifNull: ["$t", 0] }, { $multiply: [ { $ifNull: ["$v", 0] }, WEIGHT ] } ] } } },
{ $sort: { h: -1 } },
{ $limit: 20 }
]).toArray()
So sieht unsere Ausgabe aus:
{
_id: ObjectId('573a139af29313caabcf124d'),
title: 'Star Wars: Episode III - Revenge of the Sith',
plot: 'As the Clone Wars near an end, the Sith Lord Darth Sidious steps out of the shadows, at which time Anakin succumbs to his emotions, becoming Darth Vader and putting his relationships with Obi-Wan and Padme at risk.',
t: 4.957413673400879,
v: 0.756283700466156,
h: 5.2221129685640335
},
{
_id: ObjectId('573a13a6f29313caabd17d08'),
title: 'Star',
plot: 'The Driver now carries an arrogant rock star who is visiting a major city (not Pittsburgh as earlier believed). Played by Madonna, this title character wants to get away from her bodyguards...',
t: 5.151784420013428,
v: null,
h: 5.151784420013428
},
{
_id: ObjectId('573a1397f29313caabce8cdb'),
title: 'Star Wars: Episode VI - Return of the Jedi',
plot: 'After rescuing Han Solo from the palace of Jabba the Hutt, the rebels attempt to destroy the second Death Star, while Luke struggles to make Vader return from the dark side of the Force.',
t: 4.726564407348633,
v: 0.7782549858093262,
h: 4.998953652381897
},
{
_id: ObjectId('573a13aef29313caabd2da15'),
title: 'Star Runner',
plot: 'Get ready for the ultimate martial arts competition, where anything goes and lives are bought and sold. Tank is the celebrated Champion Star Runner and is deemed invincible among the ...',
t: 4.702453136444092,
v: 0.696668267250061,
h: 4.9462870299816135
},
{
_id: ObjectId('573a13c0f29313caabd62f62'),
title: 'Star Wars: The Clone Wars',
plot: 'Anakin Skywalker and Ahsoka Tano must rescue the kidnapped son of Jabba the Hutt, but political intrigue complicates their mission.',
t: 4.537533760070801,
v: 0.7561323642730713,
h: 4.802180087566375
},
{
_id: ObjectId('573a1397f29313caabce6f53'),
title: 'Message from Space',
plot: 'In this Star Wars take-off, the peaceful planet of Jillucia has been nearly wiped out by the Gavanas, whose leader takes orders from his mother (played a comic actor in drag) rather than ...',
t: 4.401333808898926,
v: 0.7913960218429565,
h: 4.678322416543961
},
{
_id: ObjectId('573a139df29313caabcfa90b'),
title: 'Message from Space',
plot: 'In this Star Wars take-off, the peaceful planet of Jillucia has been nearly wiped out by the Gavanas, whose leader takes orders from his mother (played a comic actor in drag) rather than ...',
t: 4.401333808898926,
v: 0.7913960218429565,
h: 4.678322416543961
},
{
_id: ObjectId('573a1398f29313caabce9851'),
title: 'Gymkata',
plot: 'Johnathan Cabot is a champion gymnast. In the tiny, yet savage, country of Parmistan, there is a perfect spot for a "star wars" site. For the US to get this site, they must compete in the ...',
t: 4.284864902496338,
v: 0.7193170189857483,
h: 4.53662585914135
},
{
_id: ObjectId('573a1397f29313caabce68f6'),
title: 'Star Wars: Episode IV - A New Hope',
plot: "Luke Skywalker joins forces with a Jedi Knight, a cocky pilot, a wookiee and two droids to save the universe from the Empire's world-destroying battle-station, while also attempting to rescue Princess Leia from the evil Darth Vader.",
t: 2.9629626274108887,
v: 0.7987099885940552,
h: 3.242511123418808
},
{
_id: ObjectId('573a139af29313caabcf0f5f'),
title: 'Star Wars: Episode I - The Phantom Menace',
plot: 'Two Jedi Knights escape a hostile blockade to find allies and come across a young boy who may bring balance to the Force, but the long dormant Sith resurface to reclaim their old glory.',
t: 2.9629626274108887,
v: 0.7771433591842651,
h: 3.2349628031253816
},
{
_id: ObjectId('573a1397f29313caabce77d9'),
title: 'Star Wars: Episode V - The Empire Strikes Back',
plot: 'After the rebels have been brutally overpowered by the Empire on their newly established base, Luke Skywalker takes advanced Jedi training with Master Yoda, while his friends are pursued by Darth Vader as part of his plan to capture Luke.',
t: 2.7174296379089355,
v: 0.7810181975364685,
h: 2.9907860070466996
},
{
_id: ObjectId('573a139af29313caabcf1258'),
title: 'Star Wars: Episode II - Attack of the Clones',
plot: 'Ten years after initially meeting, Anakin Skywalker shares a forbidden romance with Padmè, while Obi-Wan investigates an assassination attempt on the Senator and discovers a secret clone army crafted for the Jedi.',
t: 2.7174296379089355,
v: 0.7400004267692566,
h: 2.9764297872781755
},
{
_id: ObjectId('573a1394f29313caabcdf65b'),
title: 'Ugetsu',
plot: 'A fantastic tale of war, love, family and ambition set in the midst of the Japanese Civil Wars of the sixteenth century.',
t: 2.858368396759033,
v: null,
h: 2.858368396759033
},
{
_id: ObjectId('573a13a4f29313caabd1137f'),
title: 'S1m0ne',
plot: "A producer's film is endangered when his star walks off, so he decides to digitally create an actress to substitute for the star, becoming an overnight sensation that everyone thinks is a real person.",
t: 2.842216730117798,
v: null,
h: 2.842216730117798
},
{
_id: ObjectId('573a13b8f29313caabd4c3c3'),
title: 'Star Trek',
plot: "The brash James T. Kirk tries to live up to his father's legacy with Mr. Spock keeping him in check as a vengeful, time-traveling Romulan creates black holes to destroy the Federation one planet at a time.",
t: 2.577816963195801,
v: 0.7295459508895874,
h: 2.8331580460071564
},
{
_id: ObjectId('573a13b5f29313caabd42e99'),
title: 'Sars Wars',
plot: "The fourth generation of the virus SARS is found in Africa! It's more dangerous and causes the patients to transform into bloodthirsty zombies. The virus quickly lands to Thailand, Dr. ...",
t: 2.826810598373413,
v: null,
h: 2.826810598373413
},
{
_id: ObjectId('573a13b5f29313caabd42e1b'),
title: 'Sars Wars',
plot: "The fourth generation of the virus SARS is found in Africa! It's more dangerous and causes the patients to transform into bloodthirsty zombies. The virus quickly lands to Thailand, Dr. ...",
t: 2.826810598373413,
v: null,
h: 2.826810598373413
}
As we can see, we have some results that are identical to our query, “star wars,” and other results that are clearly off of meaning.
Conclusion
Congratulations! You have successfully completed hybrid search in MongoDB. While $rankFusion will simplify this process, this method shows a workaround while the operator is still in Public Preview. Through this tutorial, we have successfully incorporated both full-text search and vector search into one pipeline to retrieve the most optimal results for our given query. For more information on hybrid search in MongoDB, please refer to the MongoDB documentation. If you’re still getting up to speed with MongoDB, I recommend the Introduction to MongoDB in Python course.
MongoDB Hybrid Search FAQs
Do I need MongoDB 8.1 and `$rankFusion` to do hybrid search
No. On 8.0, you can actually fuse the scores yourself. Do this by running $search and $vectorSearch in two pipelines and combine them with $unionWith + $group and a weighted formula. On 8.1+, $rankFusion will do this for you.
What is hybrid search?
It’s combining full-text search (exact words) and vector search (meaning) for the most optimal results possible from a given query.
When should I use hybrid search?
Use hybrid search when your text is varied (synonyms, paraphrases, typos) but you still want exact terms.
How do I pick the weight I want to use?
It’s best practice to begin around 0.2-0.5. The lower is more for full-text search influence, and higher is for semantic influence. It’s important to tune the weight after testing and viewing the results provided.
Can hybrid search work with image and audio data?
Yes! As long as your data can be turned into vector embeddings and you can combine them with any specific wording constraints, you can perform hybrid search on a dataset.