Research any topic before you write.

Find related topics. | Discover entities. | See connections. | Build a topical map.

Word2vec

Word2vec is a technique in natural language processing for obtaining vector representations of words. These vectors capture information about the meaning of the word based on the surrounding words, following the principles of distributional semantics. Once trained, the model can be used to find words with similar meanings or usage, while its embeddings…

History & Products

Use the mouse wheel or two fingers (on touchscreens) to zoom in and out of the map.

Research this topic

Explore the main themes, entities and connections around Word2vec. Start with the topic map, then use the sections below for research and deeper semantic analysis.

Explore this topic

Start with a few of the strongest sections from the source topic. These are research directions, not a list of keywords you must use.

Key facts & relationships

High-confidence facts extracted from structured source data. Use them as anchors for further research.

License
Apache-2.0
Original author
Google AI
Release
July 29, 2013; 13 years ago (July 29, 2013)
Repository
https://code.google.com/archive/p/word2vec/
Type
Language model · Word embedding

Topics to explore

Browse the full topic structure. Each item opens a new analysis centered on that subject.

Overview

Approach

Mathematical details

History

Parameterization

Extensions

Analysis

Preservation of semantic and syntactic relationships

Assessing the quality of a model

Advanced semantic analysis

Deeper signals for content research, entity SEO and topical coverage. The plain-language headings explain what each technical view is useful for.

Map overview Semantic statistics

Word2vec

Nodes66
Edges65
Triples94
Avg. degree1.97
Density0.030303
Components1

How this topic connects Entity context

See the strongest relationship patterns around the current topic before diving into the raw triples.

Word2vec

Top relations

related to Preservation of semantic and syntactic relationships · 13
Word2vec → Brother, Capital, Country, For, Man, Mikolov, Patterns, Relationships, Sister, Such, The, This, Woman
related to doc2vec · 12
Word2vec → CBOW, Distributed Bag, Distributed Memory Model, Java, Java/Scala, Paragraph Vector, Paragraph Vectors, PV-DBOW, PV-DM, Python, The, Words
related to top2vec · 9
Word2vec → Another, As, Finally, HDBSCAN, LDA, Next, The, Together, UMAP
related to Analysis · 8
Word2vec → Arora, Firth's, Goldberg, However, Levy, The, They, Transferring
related to history · 8
Word2vec → Google, In, NeurIPS Test, Research, The, Time Award, Tomáš Mikolov, Yoshua Bengio
related to Radiology and intelligent word embeddings (IWE) · 8
Word2vec → An, Banerjee, If, IWE, Of, One, OOV, This
related to Parameters and model quality · 7
Word2vec → Accuracy, CBOW, Each, However, In, Skip-Gram, The
related to Training algorithm · 5
Word2vec → According, As, Huffman, The, To
related to Assessing the quality of a model · 4
Word2vec → Mikolov, They, This, When
related to Approach · 3
Word2vec → CBOW, In, These

Important terminology Word statistics

Use these terms to understand the vocabulary surrounding the topic, not as a checklist for keyword stuffing.

Important terminology

words word model vector displaystyle corpus embeddings skip-gram vectors used cbow similar training representations context semantic trained language models neural

Entity relationships Subject–Predicate–Object triples

SubjectPredicateObjectConfidenceSrc
Word2vecLicenseApache-2.01.00infobox
Word2vecOriginal authorGoogle AI1.00infobox
Word2vecReleaseJuly 29, 2013; 13 years ago (July 29, 2013)1.00infobox
Word2vecRepositoryhttps://code.google.com/archive/p/word2vec/1.00infobox
Word2vecTypeLanguage model1.00infobox
Word2vecTypeWord embedding1.00infobox
Word2vecis atechnique in natural language processing for obtaining vector representations of words0.90text
word2vecinstance ofThis incorporates subword information into word representations and allows vectors to be constructed for words that did not appear in the training data.Static embedding methods0.80text
fastText produce context-independent word representationsinstance ofThis incorporates subword information into word representations and allows vectors to be constructed for words that did not appear in the training data.Static embedding methods0.80text
BERTinstance ofincluding the recurrent ELMo model and transformer-based models0.80text
LDAinstance ofwhereas far away word embeddings may be considered unrelated.As opposed to other topic models0.80text
top2vec provides canonical 'distance' metrics between two topicsinstance ofwhereas far away word embeddings may be considered unrelated.As opposed to other topic models0.80text

Related concept clusters Concept neighborhoods

These clusters group vocabulary that occurs around closely connected concepts in the source material.

    Connections between topic areas Semantic bridges

    Bridges can reveal useful research angles that are easy to miss in a flat list of related terms.

    Min side: 3
    For writers, content strategists, SEOs, marketers and creators — from quick topic research to advanced semantic analysis.