🧠

Model

Pubmedbert Base Embeddings

Name: Pubmedbert Base Embeddings
Author: NeuML

by NeuML hf-model--neuml--pubmedbert-base-embeddings

Nexus Index

49.3 Top 100%

S: Semantic 50

A: Authority 0

P: Popularity 69

R: Recency 99

Q: Quality 65

Tech Context

Vital Performance

1.1M DL / 30D

0.0%

Source →

Audited 49.3 FNI Score

Tiny - Params

- Context

Hot 1.1M Downloads

Commercial APACHE License

Model Information Summary
Entity Passport
Registry ID	hf-model--neuml--pubmedbert-base-embeddings
License	Apache-2.0
Provider	huggingface

📜

Cite this model

Academic & Research Attribution

BibTeX

@misc{hf_model__neuml__pubmedbert_base_embeddings,
  author = {NeuML},
  title = {Pubmedbert Base Embeddings Model},
  year = {2026},
  howpublished = {\url{https://huggingface.co/neuml/pubmedbert-base-embeddings}},
  note = {Accessed via Free2AITools Knowledge Fortress}
}

APA Style

NeuML. (2026). Pubmedbert Base Embeddings [Model]. Free2AITools. https://huggingface.co/neuml/pubmedbert-base-embeddings

🔬Technical Deep Dive

Full Specifications [+]

Quick Commands

🤗 HF Download

huggingface-cli download neuml/pubmedbert-base-embeddings

📦 Install Lib

pip install -U transformers

⚖️ Nexus Index V2.0

Methodology Index Protocol

49.3

TOP 100% SYSTEM IMPACT

Semantic (S) 50

Authority (A) 0

Popularity (P) 69

Recency (R) 99

Quality (Q) 65

💬 Index Insight

FNI V2.0 for Pubmedbert Base Embeddings: Semantic (S:50), Authority (A:0), Popularity (P:69), Recency (R:99), Quality (Q:65).

Free2AITools Nexus Index

Verification Authority

HuggingFace API GitHub Metadata Arxiv Citation DB System Audit

Unbiased Data Node Refresh: VFS Live

---

🚀 What's Next?

📊

Find Training Datasets

Discover datasets compatible with this model

📈

Compare Benchmarks

See how this model ranks on standard tests

⚡

Technical Deep Dive

PubMedBERT Embeddings

This is a PubMedBERT-base model fined-tuned using sentence-transformers. It maps sentences & paragraphs to a 768 dimensional dense vector space and can be used for tasks like clustering or semantic search. The training dataset was generated using a random sample of PubMed title-abstract pairs along with similar title pairs.

PubMedBERT Embeddings produces higher quality embeddings than generalized models for medical literature. Further fine-tuning for a medical subdomain will result in even better performance.

Usage (txtai)

This model can be used to build embeddings databases with txtai for semantic search and/or as a knowledge source for retrieval augmented generation (RAG).

python

import txtai

embeddings = txtai.Embeddings(path="neuml/pubmedbert-base-embeddings", content=True)
embeddings.index(documents())

# Run a query
embeddings.search("query to run")

Usage (Sentence-Transformers)

Alternatively, the model can be loaded with sentence-transformers.

python

from sentence_transformers import SentenceTransformer
sentences = ["This is an example sentence", "Each sentence is converted"]

model = SentenceTransformer("neuml/pubmedbert-base-embeddings")
embeddings = model.encode(sentences)
print(embeddings)

Usage (Hugging Face Transformers)

The model can also be used directly with Transformers.

python

from transformers import AutoTokenizer, AutoModel
import torch

# Mean Pooling - Take attention mask into account for correct averaging
def meanpooling(output, mask):
    embeddings = output[0] # First element of model_output contains all token embeddings
    mask = mask.unsqueeze(-1).expand(embeddings.size()).float()
    return torch.sum(embeddings * mask, 1) / torch.clamp(mask.sum(1), min=1e-9)

# Sentences we want sentence embeddings for
sentences = ['This is an example sentence', 'Each sentence is converted']

# Load model from HuggingFace Hub
tokenizer = AutoTokenizer.from_pretrained("neuml/pubmedbert-base-embeddings")
model = AutoModel.from_pretrained("neuml/pubmedbert-base-embeddings")

# Tokenize sentences
inputs = tokenizer(sentences, padding=True, truncation=True, return_tensors='pt')

# Compute token embeddings
with torch.no_grad():
    output = model(**inputs)

# Perform pooling. In this case, mean pooling.
embeddings = meanpooling(output, inputs['attention_mask'])

print("Sentence embeddings:")
print(embeddings)

Evaluation Results

Performance of this model compared to the top base models on the MTEB leaderboard is shown below. A popular smaller model was also evaluated along with the most downloaded PubMed similarity model on the Hugging Face Hub.

The following datasets were used to evaluate model performance.

PubMed QA
- Subset: pqa_labeled, Split: train, Pair: (question, long_answer)
PubMed Subset
- Split: test, Pair: (title, text)
PubMed Summary
- Subset: pubmed, Split: validation, Pair: (article, abstract)

Evaluation results are shown below. The Pearson correlation coefficient is used as the evaluation metric.

Model	PubMed QA	PubMed Subset	PubMed Summary	Average
all-MiniLM-L6-v2	90.40	95.92	94.07	93.46
bge-base-en-v1.5	91.02	95.82	94.49	93.78
gte-base	92.97	96.90	96.24	95.37
pubmedbert-base-embeddings	93.27	97.00	96.58	95.62
S-PubMedBert-MS-MARCO	90.86	93.68	93.54	92.69

Training

The model was trained with the parameters:

DataLoader:

torch.utils.data.dataloader.DataLoader of length 20191 with parameters:

text

{'batch_size': 24, 'sampler': 'torch.utils.data.sampler.RandomSampler', 'batch_sampler': 'torch.utils.data.sampler.BatchSampler'}

Loss:

sentence_transformers.losses.MultipleNegativesRankingLoss.MultipleNegativesRankingLoss with parameters:

text

{'scale': 20.0, 'similarity_fct': 'cos_sim'}

Parameters of the fit() method:

text

{
    "epochs": 1,
    "evaluation_steps": 500,
    "evaluator": "sentence_transformers.evaluation.EmbeddingSimilarityEvaluator.EmbeddingSimilarityEvaluator",
    "max_grad_norm": 1,
    "optimizer_class": "",
    "optimizer_params": {
        "lr": 2e-05
    },
    "scheduler": "WarmupLinear",
    "steps_per_epoch": null,
    "warmup_steps": 10000,
    "weight_decay": 0.01
}

Full Model Architecture

text

SentenceTransformer(
  (0): Transformer({'max_seq_length': 512, 'do_lower_case': False}) with Transformer model: BertModel 
  (1): Pooling({'word_embedding_dimension': 768, 'pooling_mode_cls_token': False, 'pooling_mode_mean_tokens': True, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False})
)

More Information

Read more about this model and how it was built in this article.

⚠️ Incomplete Data

Some information about this model is not available. Use with Caution - Verify details from the original source before relying on this data.

View Original Source →

📝 Limitations & Considerations

• Benchmark scores may vary based on evaluation methodology and hardware configuration.
• VRAM requirements are estimates; actual usage depends on quantization and batch size.
• FNI scores are relative rankings and may change as new models are added.
⚠ License Unknown: Verify licensing terms before commercial use.

Social Proof

HuggingFace Hub

1.1MDownloads

Hub Discussions

🤗 Data Source: Hugging Face ↗

🔄 Daily sync (03:00 UTC)

AI Summary: Based on Hugging Face metadata. Not a recommendation.

📊 FNI Methodology 📚 Knowledge Baseℹ️ Verify with original source

🛡️ Model Transparency Report

Technical metadata sourced from upstream repositories.

Open Metadata

🆔 Identity & Source

id: hf-model--neuml--pubmedbert-base-embeddings
slug: neuml--pubmedbert-base-embeddings
source: huggingface
author: NeuML
license: Apache-2.0
tags: sentence-transformers, pytorch, safetensors, bert, feature-extraction, sentence-similarity, transformers, en, license:apache-2.0, text-embeddings-inference, endpoints_compatible, deploy:azure, region:us

⚙️ Technical Specs

architecture: null
params billions: null
context length: null
pipeline tag: sentence-similarity

📊 Engagement & Metrics

downloads: 1,079,031
stars: 0
forks: 0

Data indexed from public sources. Updated daily.

Welcome to Free2AI Tools!

Smart Search

FNI Score

You're All Set!

Cite this model

🔬Technical Deep Dive

Quick Commands

⚖️ Nexus Index V2.0

💬 Index Insight

Verification Authority

🚀 What's Next?

Find Training Datasets

Compare Benchmarks

Deployment Guide

Technical Deep Dive

PubMedBERT Embeddings

Usage (txtai)

Usage (Sentence-Transformers)

Usage (Hugging Face Transformers)

Evaluation Results

Training

Full Model Architecture

More Information

⚠️ Incomplete Data

📝 Limitations & Considerations

Social Proof

🛡️ Model Transparency Report

🆔 Identity & Source

⚙️ Technical Specs

📊 Engagement & Metrics