Consolidated from OpenCLIP multilabel experiments; report prepared 2026-09-29. Metrics below are percentages, except counts and labels/image. “mAP” means macro average precision, not classification accuracy.
CLIP-family models can score multiple labels independently, but these experiments do not support treating their raw scores as reliable, calibrated label decisions. The visual domain, checkpoint, prompts, and threshold calibration all matter.
- Plants: PE-Core-L improved original-prompt zero-shot mAP to 51.61%, versus 43.96% for SigLIP2-L and 30.32% for LAION-2B ViT-B/16. A supervised linear head on frozen PE-Core features reached 93.65% mAP / 89.47% micro F1.
- BigEarthNet RGB: PE-Core-L reached 33.39% zero-shot test mAP, versus 25.16% for SigLIP2-L. Its linear probe reached 69.03% mAP / 74.45% micro F1.
- NIH ChestX-ray14: BiomedCLIP loaded and worked in this checkout. It improved zero-shot ranking over the generic models, but even the selected prompt variant reached only 12.92% mAP / 64.33% macro AUROC. Linear probes improved this to approximately 19.4–19.6% mAP / 74.9–75.6% macro AUROC.
- Prompting: visual descriptions helped BigEarthNet development ranking; positive/negative comparisons helped plants and some NIH configurations but substantially hurt BigEarthNet. Six generic templates gave little improvement. There was no universally best prompt family.
The much larger probe gains suggest that frozen image embeddings contain useful task information which these text prototypes do not recover. The comparison does not isolate how much comes from label supervision, task-specific decision directions, or feature standardization. NIH remains the weakest transfer setting.
| Dataset | Labels | Probe fitting images | Calibration images | Complete held-out evaluation |
|---|---|---|---|---|
| Plant Pathology 2021 | 6 | 15,744 | 1,024 from train | 1,864 validation images |
| BigEarthNet v2 RGB | 19 | 237,871 | 122,342 validation | 119,825 test images |
| NIH ChestX-ray14 | 14 | 77,856 | 8,668 train-fold-0 | 25,596 test images |
Plant has 16,768 training images; the probe excludes the same 1,024 calibration images used for the original zero-shot threshold fitting. BigEarthNet retains its official geographic splits and uses only RGB, without multispectral bands or geographic metadata. NIH uses patient-grouped fold 0 for calibration and the remaining training folds for probe fitting. Its test and calibration populations each contain 2,797 patients, with no shared patients.
The prompt development sweep used different plant/BigEarthNet sample sizes, specified below. NIH has no separate validation split in this protocol: its training fold is development data, and the official test split is held out.
Pinned dataset revisions:
| Dataset | Hugging Face revision |
|---|---|
timm/plant-pathology-2021 |
e9c0d933b2013362ffa61eac42cc9cfd3ab1eb46 |
timm/bigearthnet-v2-rgb |
edb997737d3888abe5257c71d8e4fabc3c8ec34d |
timm/nih-chest-xray-14 |
c1bf579641b3256b4d49924436984b82bee5834d |
| Report name | OpenCLIP model | Pretrained tag |
|---|---|---|
| LAION-2B ViT-B/16 | ViT-B-16 |
laion2b_s34b_b88k |
| SigLIP2-L | ViT-L-16-SigLIP2-256 |
webli |
| PE-Core-L | PE-Core-L-14-336 |
meta |
| BiomedCLIP | hf-hub:microsoft/BiomedCLIP-PubMedBERT_256-vit_base_patch16_224 |
Hub checkpoint; omit --pretrained |
BiomedCLIP's cached model repository revision was
9f341de24bfb00180f1b847274256e9b65a3a32e. It required no compatibility fixes.
This is a practical comparison of checkpoints with different architectures,
resolutions, and training data. LAION-2B was evaluated only on plants here;
BiomedCLIP was evaluated only on NIH.
Images use deterministic native model evaluation preprocessing; grayscale images are converted to RGB. CUDA image inference uses bfloat16, followed by float32 normalization/storage. Text encoding and zero-shot scoring use float32. Class order comes from dataset metadata, and empty NIH label lists are retained. An empty list means none of the 14 findings is annotated, not that normality is proven.
For normalized image features v, each class has an independent normalized text
prototype. Each prompt embedding is normalized, the embeddings are averaged, and
the average is normalized again. There is no softmax across labels, forced top-k,
or requirement to predict at least one label.
| Scoring option | Score for class c |
|---|---|
cosine |
v @ t_positive[c] |
logit |
exp(logit_scale) * (v @ t_positive[c]) + logit_bias |
paired |
exp(logit_scale) * (v @ (t_positive[c] - t_negative[c])) |
Paired scoring retains the margin instead of applying a sigmoid, avoiding saturation before AP/AUROC calculation. It is rank-equivalent to a per-label positive/negative two-way softmax. Neither cosine nor a sigmoid applied to a CLIP/SigLIP score establishes calibrated class-presence probabilities.
- Original zero-shot ranking: frozen image/text encoders and untuned prompts; no target training labels fit the ranking function.
- Thresholded F1: frozen models with per-class thresholds selected to maximize calibration F1. These decisions use labeled calibration data.
- Prompt selection: choosing a family by labeled development mAP is another form of supervision, even though model weights remain frozen.
- Linear probe: supervised binary logistic heads fitted to training embeddings; regularization and thresholds selected on calibration data. The encoder stays frozen.
All rows use the original two positive prompts per label and cosine scoring. F1 uses the separate calibration populations above. The ranking metrics do not depend on those fitted thresholds.
| Dataset | Checkpoint | mAP | Macro AUROC | Macro F1 | Micro F1 |
|---|---|---|---|---|---|
| Plant | LAION-2B ViT-B/16 | 30.32 | 67.42 | 38.80 | 40.76 |
| Plant | SigLIP2-L | 43.96 | 74.98 | 48.28 | 47.28 |
| Plant | PE-Core-L | 51.61 | 80.67 | 53.61 | 51.33 |
| BigEarthNet RGB | SigLIP2-L | 25.16 | 65.60 | 31.65 | 42.85 |
| BigEarthNet RGB | PE-Core-L | 33.39 | 72.90 | 37.80 | 47.03 |
| NIH | SigLIP2-L | 9.70 | 56.86 | 14.41 | 23.75 |
| NIH | PE-Core-L | 8.63 | 54.44 | 13.37 | 21.78 |
| NIH | BiomedCLIP | 12.28 | 62.96 | 17.67 | 22.19 |
PE-Core-L improved on SigLIP2-L by 7.65 mAP points on plants and 8.24 points on BigEarthNet, but did not improve NIH zero-shot ranking. Better general-purpose performance did not guarantee better chest X-ray transfer.
For the earlier question specifically about full BigEarthNet validation, these are the ranking results across all 122,342 validation images. This split was used for threshold calibration, so it is separate from the held-out test table above.
| Checkpoint | Full validation mAP | Full validation macro AUROC |
|---|---|---|
| SigLIP2-L | 28.50 | 66.63 |
| PE-Core-L | 37.30 | 74.40 |
Per-class results are in the full comparison.
All six tested families are available through --prompt-strategy in
scripts/multilabel_zeroshot.py.
| CLI option | Positive ensemble per label | Default score |
|---|---|---|
baseline |
Original two domain phrases | cosine |
baseline-paired |
Original two domain phrases | paired |
domain-templates |
Six domain templates around the class name/subject | cosine |
visual-descriptions |
Three class-specific descriptions or synonyms | cosine |
baseline-plus-descriptions |
Original two phrases plus three descriptions | cosine |
descriptions-paired |
Original two phrases plus three descriptions | paired |
Every family supports all three datasets. All paired variants use the original
two negative phrases per class; descriptions-paired uses five positive texts,
not just the three descriptions. Each positive text gets equal weight in the
embedding average. Negative prototypes are averaged separately.
Examples include “an RGB satellite image containing permanent crops,” “a satellite image showing orchards and vineyards in regular rows,” and “a close-up photo of an apple leaf with bright orange spots.” NIH descriptions include label names and candidate radiographic cues. These are experimental descriptions, not expert-validated definitions or diagnostic criteria.
Baseline remains the default. The named paired options automatically use paired
scoring and reject a conflicting --score cosine or --score logit. Other
strategies allow an explicit score override, including
--prompt-strategy visual-descriptions --score paired. Such additional combinations
were not part of the six-family sweep. Custom --prompts file.json remains
available and is mutually exclusive with an explicit --prompt-strategy;
custom paired prompts require --score paired. Resolved strategy, scoring mode,
and exact texts are saved in each run's artifacts.
Candidates were written before the sweep and were not revised after inspecting its results. Within each model/dataset run, all families reused identical cached image embeddings. No held-out evaluation images were used in this sweep.
| Dataset | Development population | Fewest positives in any class |
|---|---|---|
| Plant | 2,048 of 16,768 train images | 136 |
| BigEarthNet RGB | 8,192 of 122,342 validation images | 25 |
| NIH | All 8,668 train-fold-0 images | 14 |
Plant and BigEarthNet samples were selected by the lowest SHA256 hashes of
42:image_id, across the entire development split, without label stratification.
They are not ordered prefixes and differ from the full-split populations above.
Rare classes and prevalence differences make cross-split mAP comparisons misleading.
| Prompt strategy | Plant PE-Core | BigEarthNet PE-Core | NIH PE-Core | NIH SigLIP2 | NIH BiomedCLIP |
|---|---|---|---|---|---|
baseline |
47.94 | 38.18 | 5.22 | 5.86 | 8.17 |
baseline-paired |
55.46 | 25.74 | 6.74 | 7.27 | 9.55 |
domain-templates |
48.00 | 38.11 | 4.96 | 5.71 | 8.42 |
visual-descriptions |
53.36 | 40.62 | 5.34 | 5.83 | 8.66 |
baseline-plus-descriptions |
51.78 | 40.09 | 5.30 | 5.91 | 8.76 |
descriptions-paired |
58.24 | 31.64 | 5.86 | 7.17 | 9.89 |
| Prompt strategy | Plant PE-Core | BigEarthNet PE-Core | NIH PE-Core | NIH SigLIP2 | NIH BiomedCLIP |
|---|---|---|---|---|---|
baseline |
79.71 | 74.28 | 53.63 | 54.91 | 59.96 |
baseline-paired |
80.69 | 63.64 | 61.18 | 62.39 | 67.83 |
domain-templates |
79.79 | 73.59 | 52.34 | 53.98 | 61.04 |
visual-descriptions |
81.48 | 76.71 | 52.36 | 54.34 | 60.08 |
baseline-plus-descriptions |
80.99 | 76.06 | 53.04 | 54.82 | 60.74 |
descriptions-paired |
82.03 | 71.39 | 56.41 | 61.18 | 67.16 |
Descriptions improved PE-Core BigEarthNet development mAP from 38.18 to 40.62; baseline paired scoring reduced it to 25.74. Plant's best development family was the description mixture with paired scoring, at 58.24 mAP. These are development results: the selected richer plant and BigEarthNet prompts were not subsequently evaluated on their complete held-out splits in this experiment.
All candidate prompts fit their model's context. Maximum token counts were 22/32 for plant PE-Core, 19/32 for BigEarthNet PE-Core, 21/32 for NIH PE-Core, 17/64 for NIH SigLIP2, and 18/256 for NIH BiomedCLIP. Custom prompts may exceed these limits; separate short descriptions are preferable to a long concatenation.
The visual-description approach is inspired by classification by description. Per-label positive/negative comparison also appears in CheXzero, whose model was trained on chest X-rays and reports; those results cannot be credited to prompting alone. Negation can be unreliable (NegBench). DualPrompt co-occurrence prompting was discussed as further work but was not implemented or evaluated here. This experiment is not a full reproduction of any of these methods.
BiomedCLIP was trained on broad biomedical figure-caption data, rather than being a dedicated chest-radiograph checkpoint (official model card). The single winning family by calibration mAP was fixed before test inference. The higher-calibration-AUROC baseline-paired family was not substituted after seeing test results. Only the original baseline and this selected family were tested on the 25,596 held-out images.
| BiomedCLIP prompts | Test mAP | Macro AUROC | Macro F1 | Micro F1 |
|---|---|---|---|---|
baseline |
12.28 | 62.96 | 17.67 | 22.19 |
descriptions-paired |
12.92 | 64.33 | 18.44 | 22.00 |
The selected prompts improved test mAP by 0.63 percentage point using unrounded metrics. Thresholded behavior remained weak: 13.75% micro precision, 55.02% recall, and 4.25 predicted labels/image versus 1.06 annotated. Improved ranking did not translate into improved micro F1 or reliable medical decisions.
See the BiomedCLIP report, including its selection record, exact prompts, and raw scores.
The probe uses normalized projected global image embeddings: 1,024 dimensions for PE-Core-L and 512 for BiomedCLIP. Feature standardization uses only training means and standard deviations. Independent binary logistic heads have unpenalized biases; there is no class weighting, nonlinear head, augmentation, or encoder fine-tuning.
Full-batch float64 L-BFGS minimizes
mean_binary_cross_entropy + lambda * sum(weight**2) / (2 * num_classes).
The regularization grid is [1, 0.1, 0.01, 0.001, 0.0001], selected by calibration
macro AP, with a 1,500-iteration budget. Only the selected head is evaluated on
the held-out split. Per-class thresholds maximize calibration F1.
| Dataset / encoder | Original zero-shot mAP | Probe mAP | Probe macro AUROC | Probe macro F1 | Probe micro F1 |
|---|---|---|---|---|---|
| Plant / PE-Core-L | 51.61 | 93.65 | 98.26 | 89.04 | 89.47 |
| BigEarthNet / PE-Core-L | 33.39 | 69.03 | 93.63 | 64.63 | 74.45 |
| NIH / BiomedCLIP | 12.28 | 19.39 | 74.91 | 24.85 | 33.68 |
| NIH / PE-Core-L | 8.63 | 19.61 | 75.56 | 25.15 | 32.74 |
The probe's mAP improvement over original prompts was approximately 42.04 points on plants, 35.64 on BigEarthNet, 7.11 on NIH BiomedCLIP, and 10.99 on NIH PE-Core. NIH PE-Core's weak text-based ranking did not prevent its visual features from supporting a probe comparable to BiomedCLIP. The small difference between the two NIH probes has not been tested for statistical significance.
| Dataset / encoder | Selected lambda | Micro precision | Micro recall | Predicted labels/image | Annotated labels/image |
|---|---|---|---|---|---|
| Plant / PE-Core-L | 0.01 | 89.54 | 89.41 | 1.08 | 1.08 |
| BigEarthNet / PE-Core-L | 0.001 | 70.64 | 78.69 | 3.34 | 3.00 |
| NIH / BiomedCLIP | 0.0001 | 24.68 | 52.99 | 2.28 | 1.06 |
| NIH / PE-Core-L | 0.001 | 24.01 | 51.48 | 2.28 | 1.06 |
Aggregate results hide important class differences. For example, PE-Core's BigEarthNet probe reached 99.34% AP for marine waters but 18.51% for beaches/dunes/ sands. Its NIH probe reached 41.04% AP for effusion but 4.12% for pneumonia and 3.51% for hernia. The BiomedCLIP NIH probe's calibrated hernia F1 was 0.00%, despite 78.92% AUROC. Ranking quality and a useful decision threshold are distinct issues.
All per-class AP/AUROC/F1 results, saved heads, and regularization diagnostics are in the linear-probe report.
See the usage guide for dependencies, streaming,
custom JSON format, and score definitions. Set PYTHONPATH=src to use this
checkout. --revision pins the dataset, not the model weights.
Original full BigEarthNet PE-Core evaluation, including full validation calibration:
PYTHONPATH=src python scripts/multilabel_zeroshot.py \
--dataset bigearthnet-v2-rgb \
--revision edb997737d3888abe5257c71d8e4fabc3c8ec34d \
--model PE-Core-L-14-336 --pretrained meta \
--prompt-strategy baseline --calibrate per-class \
--device cuda --amp --batch-size 128 --workers 4 --download-data \
--limit 0 --calibration-limit 0 \
--output results/multilabel/reproduce-bigearth-pecoreFor original plant results, use its dataset/revision, --calibration-limit 1024,
and --workers 0 to preserve the original ordered calibration prefix. For NIH,
the preset automatically filters calibration to patient fold 0. SigLIP2 uses
--model ViT-L-16-SigLIP2-256 --pretrained webli; the model table lists other choices.
The selected BiomedCLIP family is now directly expressible through the public script:
PYTHONPATH=src python scripts/multilabel_zeroshot.py \
--dataset nih-chest-xray-14 \
--revision c1bf579641b3256b4d49924436984b82bee5834d \
--model hf-hub:microsoft/BiomedCLIP-PubMedBERT_256-vit_base_patch16_224 \
--prompt-strategy descriptions-paired --calibrate per-class \
--device cuda --amp --batch-size 128 --workers 4 --download-data \
--output results/multilabel/reproduce-nih-biomedclip-pairedEach run saves metrics.json, exact prompts.json, and evaluation.npz; calibrated
runs also save calibration.npz. Add --save-features to retain normalized image
embeddings. Limits default to zero, meaning the complete split/fold. A positive
limit takes a prefix; it does not recreate the hash samples used by the prompt sweep.
Changing prompts requires fitting new thresholds on calibration data.
Probes can be refit from saved feature caches without rerunning the image
encoders via scripts/linear_probe.py --features <manifest.json>. A feature manifest maps
paths.train, paths.calibration, and paths.evaluation to NPZ files containing
features, targets, image_ids, patient_ids, and classnames. All caches must
use identical checkpoint/preprocessing and disjoint data partitions. The zero-shot
script saves its evaluation/calibration features; building the probe's training
cache requires a separate extraction pass with calibration images excluded.
Implementation update: scripts/linear_probe.py now
includes dataset loading/extraction, optional embedding caches, single-label and
multilabel fitting, and saved-head evaluation. It accepts HF classification
datasets, CSV, and WebDataset inputs. Both task types use this single entry point.
See the linear-probe guide and
dataset configurations for the integrated workflow.
The measurements above are unchanged; the new plant example uses a 10% hash
holdout rather than the historical 1,024-image prefix.
A saved probe predicts logits with
((features - feature_mean) / feature_std) @ weight + bias; compare those logits
to the saved per-class thresholds. The text encoder is not used by the probe.
Complete split counts, class order, label coverage, unique image IDs, identical comparison populations, and train/calibration/evaluation disjointness were checked. NIH patient IDs are also disjoint. All evaluated classes had positives and negatives, so no threshold fallback was needed. Empty multilabel targets were kept.
The logistic optimizer was compared against independent scikit-learn logistic
regression. Saved heads reconstructed evaluation logits within approximately
6e-14; saved scores reproduced the reported metrics. The six new built-in prompt
families match all 30 saved experimental prompt JSONs exactly. Automated tests
cover scoring, thresholds, split guards, prompt options, saved configuration,
custom JSON, and the probe optimizer.
Small numerical differences can occur when bfloat16 image inference uses different batch shapes. Re-extracting features for the probes retained the same displayed zero-shot mAP, but produced plant/BigEarthNet zero-shot micro F1 of 51.21/46.99, versus 51.33/47.03 in the original runs. The original table above preserves the original measurements; the linear-probe report records its own matched-cache comparison.
BigEarthNet's selected probe lambda 0.001 reached L-BFGS's default function-evaluation
budget with maximum absolute gradient 1.36e-7. The unselected lambda 0.0001
reached 1,500 iterations with maximum gradient 5.89e-8. Other candidates stopped
before their budgets. These are measured results from a fixed sweep, not a claim
of exhaustive optimization or the best attainable linear probe.
No confidence intervals, external-domain validation, or pretraining-overlap audit were performed. Prompts were evaluated on limited development populations for plants/BigEarthNet, with few positives for some classes. NIH labels are report-mined and noisy; these metrics measure agreement with those labels, not clinical reliability. Probes use projected global embeddings; intermediate features or encoder fine-tuning could give different results and were not tested here.
This report is self-contained for the main measurements. Detailed reports are included alongside it in this gist; per-run metrics, prompt JSONs, feature caches and saved heads are not: