huggingface/sentence-transformers
State-of-the-Art Embeddings, Retrieval, and Reranking
https://skillcdn.ai/gh/huggingface/sentence-transformersConnect this address to your AI to use these skills. How to connect
- Unverified
- Default branch
main - Commit
4a3b5cd - License
Apache-2.0
Explore this repository
Searches the original text of skills and documents; display translations are not searched.
Documents
- MS MARCO Cross-EncodersMS MARCO is a large scale information retrieval corpus that was created based on real user search queries using Bing search engine. The provided models can be used for semantic search, i.e., given ke…docs/pretrained-models/ce-msmarco.md
- DPR-ModelsIn Dense Passage Retrieval for Open-Domain Question Answering Karpukhin et al. trained models based on Google's Natural Questions dataset:docs/pretrained-models/dpr.md
- MSMARCO ModelsMS MARCO is a large scale information retrieval corpus that was created based on real user search queries using Bing search engine. The provided models can be used for semantic search, i.e., given ke…docs/pretrained-models/msmarco-v1.md
- MSMARCO Models (Version 2)MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using Bing search engine. The provided models can be used for semantic search, i.e., given ke…docs/pretrained-models/msmarco-v2.md
- MSMARCO ModelsMS MARCO is a large scale information retrieval corpus that was created based on real user search queries using Bing search engine. The provided models can be used for semantic search, i.e., given ke…docs/pretrained-models/msmarco-v3.md
- MSMARCO ModelsMS MARCO is a large scale information retrieval corpus that was created based on real user search queries using Bing search engine. The provided models can be used for semantic search, i.e., given ke…docs/pretrained-models/msmarco-v5.md
- NLI ModelsConneau et al., 2017, show in the InferSent-Paper (Supervised Learning of Universal Sentence Representations from Natural Language Inference Data) that training on Natural Language Inference (NLI) da…docs/pretrained-models/nli-models.md
- Natural Questions ModelsGoogle's Natural Questions dataset consists of about 100k real search queries from Google with the respective, relevant passage from Wikipedia. Models trained on this dataset work well for question-a…docs/pretrained-models/nq-v1.md
- STS ModelsThe models were first trained on NLI data, then we fine-tuned them on the STS benchmark dataset (docs, dataset). This generate sentence embeddings that are especially suitable to measure the semantic…docs/pretrained-models/sts-models.md
- Wikipedia Sections ModelsThe wikipedia-sections-models implement the idea from Ein Dor et al., 2018, Learning Thematic Similarity Metric Using Triplet Networks.docs/pretrained-models/wikipedia-sections-models.md