<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.0 20040830//EN" "http://dtd.nlm.nih.gov/publishing/2.0/journalpublishing.dtd">
<article article-type="research-article" dtd-version="2.0" xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-id journal-id-type="publisher-id">JMI</journal-id>
      <journal-id journal-id-type="nlm-ta">JMIR Med Inform</journal-id>
      <journal-title>JMIR Medical Informatics</journal-title>
      <issn pub-type="epub">2291-9694</issn>
      <publisher>
        <publisher-name>JMIR Publications</publisher-name>
        <publisher-loc>Toronto, Canada</publisher-loc>
      </publisher>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="publisher-id">v14i1e99639</article-id>
      <article-id pub-id-type="pmid">42455615</article-id>
      <article-id pub-id-type="doi">10.2196/99639</article-id>
      <article-categories>
        <subj-group subj-group-type="heading">
          <subject>Original Paper</subject>
        </subj-group>
        <subj-group subj-group-type="article-type">
          <subject>Original Paper</subject>
        </subj-group>
      </article-categories>
      <title-group>
        <article-title>Effects of Model Choice, Corpus Context, and Post Hoc Correction on Layer-Level Embedding Degradation in Clinical Document Retrieval: Experimental Study</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="editor">
          <name>
            <surname>Coristine</surname>
            <given-names>Andrew</given-names>
          </name>
        </contrib>
      </contrib-group>
      <contrib-group>
        <contrib contrib-type="reviewer">
          <name>
            <surname>Asare</surname>
            <given-names>Ebenezer Aquisman</given-names>
          </name>
        </contrib>
        <contrib contrib-type="reviewer">
          <name>
            <surname>Elmitwalli</surname>
            <given-names>Sherif</given-names>
          </name>
        </contrib>
        <contrib contrib-type="reviewer">
          <name>
            <surname>Ma</surname>
            <given-names>Chunwei</given-names>
          </name>
        </contrib>
      </contrib-group>
      <contrib-group>
        <contrib id="contrib1" contrib-type="author" corresp="yes">
          <name name-style="western">
            <surname>Mikkelsen</surname>
            <given-names>Yngve</given-names>
          </name>
          <degrees>MD, MSc, DBA</degrees>
          <xref rid="aff1" ref-type="aff">1</xref>
          <address>
            <institution>Saïd Business School</institution>
            <institution>University of Oxford</institution>
            <addr-line>Park End Street</addr-line>
            <addr-line>Oxford, England, OX1 1HP</addr-line>
            <country>United Kingdom</country>
            <phone>44 1865 270000</phone>
            <email>yngve.mikkelsen@sbs.ox.ac.uk</email>
          </address>
          <ext-link ext-link-type="orcid">https://orcid.org/0000-0003-1543-3805</ext-link>
        </contrib>
      </contrib-group>
      <aff id="aff1">
        <label>1</label>
        <institution>Saïd Business School</institution>
        <institution>University of Oxford</institution>
        <addr-line>Oxford, England</addr-line>
        <country>United Kingdom</country>
      </aff>
      <author-notes>
        <corresp>Corresponding Author: Yngve Mikkelsen <email>yngve.mikkelsen@sbs.ox.ac.uk</email></corresp>
      </author-notes>
      <pub-date pub-type="collection">
        <year>2026</year>
      </pub-date>
      <pub-date pub-type="epub">
        <day>29</day>
        <month>7</month>
        <year>2026</year>
      </pub-date>
      <volume>14</volume>
      <elocation-id>e99639</elocation-id>
      <history>
        <date date-type="received">
          <day>27</day>
          <month>4</month>
          <year>2026</year>
        </date>
        <date date-type="rev-request">
          <day>25</day>
          <month>5</month>
          <year>2026</year>
        </date>
        <date date-type="accepted">
          <day>15</day>
          <month>7</month>
          <year>2026</year>
        </date>
      </history>
      <copyright-statement>©Yngve Mikkelsen. Originally published in JMIR Medical Informatics (https://medinform.jmir.org), 29.07.2026.</copyright-statement>
      <copyright-year>2026</copyright-year>
      <license license-type="open-access" xlink:href="https://creativecommons.org/licenses/by/4.0/">
        <p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (https://creativecommons.org/licenses/by/4.0/), which permits unrestricted use, distribution, and reproduction in any medium, provided the original work, first published in JMIR Medical Informatics, is properly cited. The complete bibliographic information, a link to the original publication on https://medinform.jmir.org/, as well as this copyright and license information must be included.</p>
      </license>
      <self-uri xlink:href="https://medinform.jmir.org/2026/1/e99639" xlink:type="simple"/>
      <abstract>
        <sec sec-type="background">
          <title>Background</title>
          <p>Clinical retrieval-augmented generation depends on embedding models. A companion study found that non–retrieval-trained encoders underperformed retrieval-trained general-purpose embeddings and produced near-degenerate embedding geometry, but did not localize the architectural origin, separate training-domain from training-objective effects, or test whether the degradation can be corrected without retraining.</p>
        </sec>
        <sec sec-type="objective">
          <title>Objective</title>
          <p>This study aimed to (1) characterize layer-wise retrieval and geometric trajectories across 13 transformer configurations on clinical documents, (2) separate training-objective from training-domain effects through matched architectural comparisons, (3) reanalyze the panel under a per-query layer-wise linear mixed-effects (LME) framework, and (4) evaluate deployment-relevant post hoc geometric correction.</p>
        </sec>
        <sec sec-type="methods">
          <title>Methods</title>
          <p>Layer-wise embeddings were extracted from 13 transformer configurations on 3 clinical corpora (n=100 documents each: MTSamples, PMC-Patients, and Mistral-7B-Instruct–generated synthetic notes) under 2 query formats (keyword and natural-language via GPT-4o). Retrieval performance (mean reciprocal rank at cutoff 10 [MRR@10] and recall at cutoff 10 [recall@10]) and geometric properties (participation ratio, average pairwise cosine, and anisotropy) were measured at every layer. A per-query layer-wise LME model was fit independently per configuration. Corpus-only zero-phase component analysis (ZCA) whitening was evaluated as the primary deployment-relevant intervention, with 5-fold cross-validation, an epsilon sweep, a lexical-overlap audit, and a chunking sensitivity analysis.</p>
        </sec>
        <sec sec-type="results">
          <title>Results</title>
          <p>Document embeddings clustered into 3 anisotropy tiers: extreme (average pairwise cosine &gt;0.92) for non–retrieval-trained encoders and large language models (LLMs), moderate (0.65-0.92) for general retrievers and most LLMs, and reduced (&lt;0.65) for BioLORD-2023, instruction-tuned E5-Mistral-7B, and Nomic-embed-text-nopfx. The per-query random-slope LME identified 2 layer-depth patterns: classical degradation with depth in 3 non–retrieval-trained encoders (all <italic>P</italic>&lt;.001), vs net improvement with depth in the remaining 10 models (all <italic>P</italic>&lt;.001), with the strongest negative coefficients in decoder LLMs. Matched-contrast tests confirmed significant training-objective × layer-depth interactions in all 3 matched pairs (all <italic>P</italic>&lt;.001). Corpus-only ZCA whitening produced a 2-tier pattern under 5-fold cross-validation: tier 2 non–retrieval-trained models showed positive ΔMRR@10 (+0.066 to +0.304), while tier 1 retrieval-trained models showed negative ΔMRR@10 (–0.021 to –0.051). Best Match 25-vs-embedding Spearman rank correlations spanned –0.02 to 0.37, indicating substantial nonlexical contribution to retrieval. Same-source ranking stability replicated at 4-5× corpus scale (ρ=0.952 for PMC-500, ρ=0.929 for MTSamples-400) for the BERT-scale subset.</p>
        </sec>
        <sec sec-type="conclusions">
          <title>Conclusions</title>
          <p>Anisotropy in transformer embeddings is widespread across architectural classes and is lower in configurations with retrieval-specific training. Corpus-only ZCA whitening is a deployment-compatible, retraining-free post hoc correction candidate that improved retrieval for non–retrieval-trained models on this controlled benchmark but requires target-corpus validation before clinical deployment. The matched-comparison evidence supports training objective rather than training domain as the stronger explanatory axis, though residual confounding is not eliminated. The principal contribution is mechanistic: layer-level localization of embedding degradation and the geometric basis for the 2-tier intervention response.</p>
        </sec>
      </abstract>
      <kwd-group>
        <kwd>anisotropy</kwd>
        <kwd>clinical decision support</kwd>
        <kwd>clinical informatics</kwd>
        <kwd>embedding geometry</kwd>
        <kwd>retrieval-augmented generation</kwd>
        <kwd>transformer layers</kwd>
        <kwd>ZCA whitening</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec sec-type="introduction">
      <title>Introduction</title>
      <sec>
        <title>Overview</title>
        <p>Retrieval-augmented generation (RAG) is the dominant approach for grounding large language model (LLM) outputs in verifiable evidence sources [<xref ref-type="bibr" rid="ref1">1</xref>,<xref ref-type="bibr" rid="ref2">2</xref>]. In a clinical RAG pipeline, the embedding model transforms queries and documents into vector representations, and retrieval quality depends on the geometry of the resulting embedding space. When embeddings cluster in a narrow region of the vector space, a phenomenon known as anisotropy, cosine similarity becomes less discriminative, and retrieval may fail regardless of downstream generation quality.</p>
        <p>Retrieval failure in clinical RAG settings may affect downstream outputs consumed by practitioners or clinical decision support systems. This study does not evaluate clinical task performance, patient outcomes, or production electronic health record (EHR) deployment; it addresses the upstream question of where in the model architecture retrieval performance is determined and whether geometric properties can be corrected without retraining.</p>
      </sec>
      <sec>
        <title>Companion Study and Motivating Context</title>
        <p>A recently accepted companion benchmark study [<xref ref-type="bibr" rid="ref3">3</xref>] evaluated 10 embedding models across 3 clinical corpora and 294 experimental conditions. Two findings motivated the present analysis. First, model choice and clinical-context variables (corpus type and query format) contributed comparably to variance in mean reciprocal rank at cutoff 10 (MRR@10). Second, domain-specific models, BioBERT [<xref ref-type="bibr" rid="ref4">4</xref>] and ClinicalBERT [<xref ref-type="bibr" rid="ref5">5</xref>], produced near-degenerate embedding geometry, with pairwise cosine similarities exceeding 0.90, whereas general-purpose models (Beijing Academy of Artificial Intelligence General Embedding [BGE]-base [<xref ref-type="bibr" rid="ref6">6</xref>], General Text Embeddings [GTE]-base [<xref ref-type="bibr" rid="ref7">7</xref>], Nomic-embed-text [<xref ref-type="bibr" rid="ref8">8</xref>]) achieved higher MRR@10 and more dispersed geometry. The companion study established that domain-pretrained encoders underperform but did not localize <italic>where</italic> in the network the gap appears, <italic>what</italic> causes it, or <italic>whether</italic> it can be corrected.</p>
        <p>Comparing BERT-architecture domain encoders with retrieval-trained text-embedding models can conflate training domain (biomedical vs general) with training objective (retrieval-trained vs not). The present analysis therefore adds matched architectural controls to separate these dimensions as much as possible within the public model space.</p>
      </sec>
      <sec>
        <title>Scope and Demarcation</title>
        <p>Retrieval failures in production systems arise from multiple sources beyond embedding model behavior: corpus heterogeneity, query formulation, chunking strategy, indexing choices, and vocabulary mismatch between queries and documents. This study isolates the embedding model’s contribution by holding corpus, query generation, and indexing constant across conditions, while the factorial design models residual corpus and query-format effects directly. Chunking effects were negligible in the companion study (η<sup>2</sup>=0.002) [<xref ref-type="bibr" rid="ref3">3</xref>] but are reassessed here on a representative subset of 4 models to ensure robustness. The interventions evaluated here target embedding geometry; they do not address retrieval failures driven by vocabulary mismatch or chunking strategy at scale.</p>
      </sec>
      <sec>
        <title>Research Questions</title>
        <p>The four research questions addressed are:</p>
        <list list-type="bullet">
          <list-item>
            <p>RQ1: How does retrieval performance vary with transformer layer depth across the 13 model configurations, and what layer-wise patterns are consistent across architectural classes?</p>
          </list-item>
          <list-item>
            <p>RQ2: How much variation in retrieval performance is associated with model configuration vs clinical-context factors (corpus, query format), and does this attribution change with layer depth?</p>
          </list-item>
          <list-item>
            <p>RQ3: In matched architectural comparisons that hold the training objective constant or the training domain constant, which axis, the training objective or the training domain, drives the performance gap between models?</p>
          </list-item>
          <list-item>
            <p>RQ4: Can deployment-relevant post hoc geometric correction (corpus-only zero-phase component analysis [ZCA] whitening, fit on document embeddings alone) recover retrieval performance to a substantively meaningful extent in this controlled benchmark, and does this correction generalize across documents not used to fit the transform?</p>
          </list-item>
        </list>
        <p>RQ2 is phrased in terms of association rather than causation because the study cannot fully isolate training objective from architecture, pooling, model size, tokenizer, and training data; it therefore quantifies variation associated with model configuration as a whole. RQ3 addresses the training-objective vs training-domain contrast using matched architectural comparisons.</p>
        <p>The intended contribution is mechanistic and methodological. This study does not propose a deployment standard; practical implications are presented as discussion-level diagnostic considerations that require prospective validation on target institutional corpora before use in clinical retrieval systems.</p>
      </sec>
      <sec>
        <title>Relationship to Companion Study</title>
        <p>This manuscript is a mechanistic companion to the accepted benchmark study [<xref ref-type="bibr" rid="ref3">3</xref>]. It reuses the same clinical retrieval setting, corpora, and metadata-derived query design, but adds layer-wise extraction, geometric trajectory analysis, matched controls, ZCA evaluation, and per-query, layer-wise modeling. The 2 manuscripts are therefore complementary rather than duplicative.</p>
      </sec>
      <sec>
        <title>Related Work</title>
        <sec>
          <title>Anisotropy and Embedding Geometry in Transformer Representations</title>
          <p>The anisotropy of contextual word representations was characterized by Ethayarajh [<xref ref-type="bibr" rid="ref9">9</xref>], who showed that contextual embeddings from BERT, GPT-2, and ELMo occupy a narrow cone in representation space across all noninput layers. Razzhigaev et al [<xref ref-type="bibr" rid="ref10">10</xref>] demonstrated distinct anisotropy profiles for encoder vs decoder models. Mechanistic accounts diverge: Machina and Mercer [<xref ref-type="bibr" rid="ref11">11</xref>] attribute anisotropy to tied input/output embedding matrices, whereas Godey et al [<xref ref-type="bibr" rid="ref12">12</xref>] link it to sharp attention patterns. Subsequent work has explored isotropy-promoting training objectives (BioLORD-2023 [<xref ref-type="bibr" rid="ref13">13</xref>] adopts an explicit isotropy term in its contrastive loss) and post hoc whitening transforms [<xref ref-type="bibr" rid="ref14">14</xref>-<xref ref-type="bibr" rid="ref16">16</xref>] that compress eigenspectra. These studies focus on general-domain text; to the author’s knowledge, no prior work has conducted layer-level geometric and retrieval analyses of clinical or biomedical text across the architectural classes evaluated here.</p>
        </sec>
        <sec>
          <title>Clinical Embedding Benchmarks</title>
          <p>In the clinical embedding benchmarking literature, Myers et al [<xref ref-type="bibr" rid="ref17">17</xref>] evaluated 7 embedding models with different pooling strategies for EHR retrieval and reported that BGE-base outperformed domain-specific models, consistent with the companion study’s findings. Soffer et al [<xref ref-type="bibr" rid="ref18">18</xref>] evaluated 39 embedding models across 7 medical semantic similarity tasks. Hosseini et al [<xref ref-type="bibr" rid="ref19">19</xref>] showed that combining intermediate BERT layers improved semantic textual similarity by up to 16% in Spearman correlation, though this was evaluated on general-domain similarity tasks rather than clinical retrieval. None of these studies conducted layer-level analysis on clinical retrieval data.</p>
        </sec>
        <sec>
          <title>Retrieval Evaluation Methodology</title>
          <p>The information retrieval evaluation literature provides the methodological foundation for the metrics used here. The Massive Text Embedding Benchmark [<xref ref-type="bibr" rid="ref20">20</xref>] established standardized retrieval evaluation across diverse domains, and the Benchmarking Information Retrieval (BEIR) suit [<xref ref-type="bibr" rid="ref21">21</xref>] provided heterogeneous-domain retrieval benchmarks emphasizing zero-shot generalization. Both highlight that ranking metrics computed on small candidate sets can overestimate performance relative to production-scale retrieval, a consideration that motivated the validation experiments at 4-5× scale in this study.</p>
          <p>Dense retrieval architectures pioneered by Karpukhin et al [<xref ref-type="bibr" rid="ref22">22</xref>] have been widely adopted for clinical and biomedical retrieval, with their failure modes subsequently characterized in BEIR-style evaluations [<xref ref-type="bibr" rid="ref23">23</xref>]. Reranking pipelines such as ColBERT [<xref ref-type="bibr" rid="ref24">24</xref>] and BERT-based rerankers [<xref ref-type="bibr" rid="ref25">25</xref>] partially mitigate the limitations of single-stage dense retrieval. Calibration of embedding similarity scores [<xref ref-type="bibr" rid="ref26">26</xref>] is an active area; the present ZCA correction is a form of post hoc calibration applied at the embedding-vector level rather than at the score level. Clinical information retrieval evaluation has emphasized known-item retrieval [<xref ref-type="bibr" rid="ref27">27</xref>] and the impact of query-document lexical overlap on apparent performance [<xref ref-type="bibr" rid="ref23">23</xref>]; both inform the lexical-overlap audit reported below.</p>
        </sec>
        <sec>
          <title>Gap Addressed by This Study</title>
          <p>To the author’s knowledge, no prior study has conducted (1) layer-level geometric and retrieval analyses of clinical documents across 13 transformer configurations, including matched architectural controls that separate training-objective from training-domain effects; (2) layer-wise per-query linear mixed-effects (LME) analysis of retrieval performance; or (3) cross-validated evaluation of corpus-only ZCA whitening as a deployment-relevant geometric correction. This study addresses each of these gaps directly.</p>
        </sec>
      </sec>
    </sec>
    <sec sec-type="methods">
      <title>Methods</title>
      <sec>
        <title>Models</title>
        <sec>
          <title>Overview</title>
          <p>A total of 11 transformer models across 13 configurations were evaluated (<xref ref-type="table" rid="table1">Table 1</xref>). The panel aligns with the companion study [<xref ref-type="bibr" rid="ref3">3</xref>] by excluding the OpenAI application programming interface model (no hidden-state access) and Best Match 25 (BM25; nonneural), and by adding 2 new configurations to enable matched comparisons across training objective and training domain dimensions. The models span 6 categories: domain encoders (BioBERT [<xref ref-type="bibr" rid="ref4">4</xref>], ClinicalBERT [<xref ref-type="bibr" rid="ref5">5</xref>]), general encoders (BERT-base-uncased [<xref ref-type="bibr" rid="ref28">28</xref>]), biomedical retrievers (BioLORD-2023 [<xref ref-type="bibr" rid="ref13">13</xref>], MedCPT [<xref ref-type="bibr" rid="ref29">29</xref>]), general-purpose retrievers (BGE-base [<xref ref-type="bibr" rid="ref6">6</xref>], GTE-base [<xref ref-type="bibr" rid="ref7">7</xref>], Nomic-embed-text [<xref ref-type="bibr" rid="ref8">8</xref>]), general LLMs (E5-Mistral-7B [<xref ref-type="bibr" rid="ref30">30</xref>], Phi-3-mini [<xref ref-type="bibr" rid="ref31">31</xref>]), and biomedical LLMs (BioMistral-7B [<xref ref-type="bibr" rid="ref32">32</xref>]). Two ablation variants are retained from the companion design: E5-Mistral-7B with mean pooling and no instruction prefix and Nomic-embed-text without search_query/search_document prefixes.</p>
          <p>ClinicalBERT uses a DistilBERT architecture (6 transformer layers) rather than BERT-base (12 layers); this architectural difference is noted where relevant. BERT-base-scale models have 13 extraction points (input embeddings plus 12 transformer layers). LLM-scale models (E5-Mistral-7B, Phi-3-mini, BioMistral-7B, and the E5-Mistral-7B-ablation variant) have 33 extraction points (including the input and 32 layers). MedCPT, a dual-encoder model, was evaluated by extracting layers from the query and article encoders separately.</p>
          <table-wrap position="float" id="table1">
            <label>Table 1</label>
            <caption>
              <p>Model configurations evaluated in the 13-configuration panel. BERT-base-uncased and BioMistral-7B were added beyond the companion-study panel to serve as matched-comparison controls.</p>
            </caption>
            <table width="1000" cellpadding="5" cellspacing="0" border="1" rules="groups" frame="hsides">
              <col width="150"/>
              <col width="260"/>
              <col width="110"/>
              <col width="140"/>
              <col width="80"/>
              <col width="80"/>
              <col width="180"/>
              <thead>
                <tr valign="top">
                  <td>Model</td>
                  <td>HuggingFace ID</td>
                  <td>Category<sup>a</sup></td>
                  <td>Architecture<sup>b</sup></td>
                  <td>Layers</td>
                  <td>Pool<sup>c</sup></td>
                  <td>Notes</td>
                </tr>
              </thead>
              <tbody>
                <tr valign="top">
                  <td>BioBERT</td>
                  <td>dmis-lab/biobert-v1.1</td>
                  <td>DE<sup>d</sup></td>
                  <td>BERT-base</td>
                  <td>12</td>
                  <td>M</td>
                  <td>fp32<sup>e</sup></td>
                </tr>
                <tr valign="top">
                  <td>ClinicalBERT</td>
                  <td>medicalai/ClinicalBERT</td>
                  <td>DE</td>
                  <td>DistilBERT</td>
                  <td>6</td>
                  <td>M</td>
                  <td>fp32</td>
                </tr>
                <tr valign="top">
                  <td>BERT-base</td>
                  <td>bert-base-uncased</td>
                  <td>GE<sup>f</sup></td>
                  <td>BERT-base</td>
                  <td>12</td>
                  <td>M</td>
                  <td>fp32; matched control</td>
                </tr>
                <tr valign="top">
                  <td>BioLORD</td>
                  <td>FremyCompany/BioLORD</td>
                  <td>BR<sup>g</sup></td>
                  <td>MPNet</td>
                  <td>12</td>
                  <td>M</td>
                  <td>fp32</td>
                </tr>
                <tr valign="top">
                  <td>MedCPT</td>
                  <td>ncbi/MedCPT-{Query,Article}-Encoder</td>
                  <td>BR</td>
                  <td>BERT-base</td>
                  <td>12</td>
                  <td>C</td>
                  <td>Dual encoder</td>
                </tr>
                <tr valign="top">
                  <td>BGE<sup>h</sup>-base</td>
                  <td>BAAI/bge-base-en-v1.5</td>
                  <td>GR<sup>i</sup></td>
                  <td>BERT-base</td>
                  <td>12</td>
                  <td>M</td>
                  <td>fp32</td>
                </tr>
                <tr valign="top">
                  <td>GTE<sup>j</sup>-base</td>
                  <td>thenlper/gte-base</td>
                  <td>GR</td>
                  <td>BERT-base</td>
                  <td>12</td>
                  <td>M</td>
                  <td>fp32</td>
                </tr>
                <tr valign="top">
                  <td>Nomic</td>
                  <td>nomic-ai/nomic-embed-text-v1.5</td>
                  <td>GR</td>
                  <td>NomicBERT</td>
                  <td>12</td>
                  <td>M</td>
                  <td>With prefixes</td>
                </tr>
                <tr valign="top">
                  <td>Nomic-nopfx</td>
                  <td>nomic-ai/nomic-embed-text-v1.5</td>
                  <td>GR</td>
                  <td>NomicBERT</td>
                  <td>12</td>
                  <td>M</td>
                  <td>Without prefixes (ablation)</td>
                </tr>
                <tr valign="top">
                  <td>E5-Mistral</td>
                  <td>intfloat/e5-mistral-7b-instruct</td>
                  <td>GL<sup>k</sup></td>
                  <td>Mistral-7B</td>
                  <td>32</td>
                  <td>E</td>
                  <td>fp16<sup>l</sup>; instruction-tuned</td>
                </tr>
                <tr valign="top">
                  <td>E5-Mistral- ablation</td>
                  <td>intfloat/e5-mistral-7b-instruct</td>
                  <td>GL</td>
                  <td>Mistral-7B</td>
                  <td>32</td>
                  <td>M</td>
                  <td>Mean pool, no instruction (ablation)</td>
                </tr>
                <tr valign="top">
                  <td>Phi-3-mini</td>
                  <td>microsoft/Phi-3-mini-4k-instruct</td>
                  <td>GL</td>
                  <td>Phi-3</td>
                  <td>32</td>
                  <td>M</td>
                  <td>fp16</td>
                </tr>
                <tr valign="top">
                  <td>BioMistral</td>
                  <td>BioMistral/BioMistral</td>
                  <td>DL<sup>m</sup></td>
                  <td>Mistral-7B</td>
                  <td>32</td>
                  <td>M</td>
                  <td>fp16; matched control</td>
                </tr>
              </tbody>
            </table>
            <table-wrap-foot>
              <fn id="table1fn1">
                <p><sup>a</sup>Model category (general encoder, biomedical encoder, general retriever, biomedical retriever, decoder LLM).</p>
              </fn>
              <fn id="table1fn2">
                <p><sup>b</sup>Base architecture.</p>
              </fn>
              <fn id="table1fn3">
                <p><sup>c</sup>Pooling strategy (CLS=classification [CLS] token; Mean=mean of token embeddings; Last=last token; EOS=end-of-sequence token embedding).</p>
              </fn>
              <fn id="table1fn4">
                <p><sup>d</sup>DE: domain encoder.</p>
              </fn>
              <fn id="table1fn5">
                <p><sup>e</sup>fp32: 32-bit floating-point precision.</p>
              </fn>
              <fn id="table1fn6">
                <p><sup>f</sup>GE: general encoder.</p>
              </fn>
              <fn id="table1fn7">
                <p><sup>g</sup>BR: biomedical retriever.</p>
              </fn>
              <fn id="table1fn8">
                <p><sup>h</sup>BGE: Beijing Academy of Artificial Intelligence General Embedding.</p>
              </fn>
              <fn id="table1fn9">
                <p><sup>i</sup>GR: general retriever.</p>
              </fn>
              <fn id="table1fn10">
                <p><sup>j</sup>GTE: General Text Embeddings.</p>
              </fn>
              <fn id="table1fn11">
                <p><sup>k</sup>GL: general large language model.</p>
              </fn>
              <fn id="table1fn12">
                <p><sup>l</sup>fp16: 16-bit floating-point precision.</p>
              </fn>
              <fn id="table1fn13">
                <p><sup>m</sup>DL: domain large language model.</p>
              </fn>
            </table-wrap-foot>
          </table-wrap>
        </sec>
        <sec>
          <title>Matched-Comparison Rationale</title>
          <p>Two configurations were added beyond the companion panel to help separate the training objective from the training domain. BERT-base-uncased serves as a non–retrieval-trained BERT-base control for BioBERT and the BERT-family retrievers. BioMistral-7B serves as a biomedical Mistral-7B control for E5-Mistral-7B. These comparisons are architecture-matched or partially matched, but residual confounding from the tokenizer, instruction data, pooling, and model scale remains.</p>
        </sec>
      </sec>
      <sec>
        <title>Corpora and Queries</title>
        <sec>
          <title>Overview</title>
          <p>Three clinical corpora were reused from the companion study [<xref ref-type="bibr" rid="ref3">3</xref>] at a 100-document working scale per corpus to enable mechanistic layer-by-layer analysis: MTSamples (n=100), PMC-Patients case reports from PubMed Central (n=100) [<xref ref-type="bibr" rid="ref20">20</xref>], and synthetic clinical notes generated by Mistral-7B-Instruct-v0.2 (n=100). For each corpus, 100 keywords and 100 natural-language queries were generated from document metadata using GPT-4o, following the reduced-lexical-dependence design used in the companion study’s validation experiment [<xref ref-type="bibr" rid="ref3">3</xref>]. Each query corresponds to a single known-relevant document (known-item retrieval design). The complete GPT-4o prompt is provided in <xref ref-type="supplementary-material" rid="app1">Multimedia Appendix 1</xref>.</p>
          <p>The known-item retrieval design with 100 documents per corpus provides a controlled environment for mechanistic, layer-by-layer analysis but is substantially smaller than production indices; absolute results should not be extrapolated directly to production-scale retrieval. Validation at 4-5× scale (500 PMC-Patients and 400 MTSamples) is reported in the “Validation on Independent Corpora” subsection of the Results section.</p>
        </sec>
        <sec>
          <title>Alignment Audit for the Synthetic Corpus</title>
          <p>The layer-extraction pipeline preserved the document order used to generate the metadata queries. A locally cached synthetic CSV used to reload the corpus for a sensitivity analysis had been resorted by synth_id, breaking the index alignment; query-document pairing for that sensitivity analysis was therefore recovered via BM25 top-1 matching with greedy one-to-one deduplication, replicating the alignment-recovery procedure documented in the companion study [<xref ref-type="bibr" rid="ref3">3</xref>]. The full audit table is reproduced in <xref ref-type="supplementary-material" rid="app2">Multimedia Appendix 2</xref>. The primary per-query layer-wise LME uses ranks derived from the original-pipeline alignment, not the recovered alignment.</p>
        </sec>
        <sec>
          <title>BM25 Alignment-Check Baselines</title>
          <p>Alignment between queries and documents was verified using BM25 on the 100-document subset before evaluating each embedding model, yielding MRR@10 of 0.87 (MTSamples), 0.92 (PMC-Patients), and 0.96 (synthetic). These baselines confirm that the queries are appropriately aligned with their target documents at the lexical level and serve as a sanity check on the experimental design. They are not evidence that retrieval can be reduced to lexical matching: the lexical-overlap audit measures the BM25-vs-embedding rank correlation directly across the 13 models and confirms a substantial nonlexical contribution to embedding-based retrieval (Spearman ρ range –0.02 to 0.37). This finding directly addresses the apparent contradiction between high BM25 baselines and the manuscript’s claim of reduced lexical dependence.</p>
        </sec>
      </sec>
      <sec>
        <title>Layer-Level Measurements</title>
        <sec>
          <title>Overview</title>
          <p>For each model and layer, embeddings were extracted across the 6 corpus × query format conditions using output_hidden_states=True (or equivalent forward hooks for architectures that do not support this parameter). Pooling matched vendor specifications: mean pooling for most models, classification token pooling for MedCPT, and last-token (end-of-sequence [EOS]) pooling for E5-Mistral-7B. All embeddings were L2-normalized.</p>
        </sec>
        <sec>
          <title>Geometric Measures (Computed on Document Embeddings)</title>
          <p>The following geometric measures were used:</p>
          <list list-type="bullet">
            <list-item>
              <p>Participation ratio (PR): PR = (sum_i sigma_i<sup>2</sup>)<sup>2</sup>/sum_i sigma_i<sup>4</sup>, where sigma_i are the singular values of the centered embedding matrix. PR reflects the effective number of dimensions used by the embedding distribution; lower values indicate more collapsed representations.</p>
            </list-item>
            <list-item>
              <p>Average pairwise cosine similarity (average pairwise cosine): the mean cosine similarity across 10,000 randomly sampled embedding pairs. Higher values indicate a narrower embedding cone (more anisotropic).</p>
            </list-item>
            <list-item>
              <p>Anisotropy singular value decomposition–based: the top squared singular value as a fraction of the sum of all squared singular values, indicating concentration of variance along a single direction.</p>
            </list-item>
            <list-item>
              <p>Retrieval measures: MRR@10 was the primary retrieval metric. Recall at cutoff 10 (Recall@10; the proportion of queries for which the relevant document appeared in the top 10 results) was reported as a secondary metric. Both were computed by ranking documents using cosine similarity with a one-to-one query-document mapping.</p>
            </list-item>
            <list-item>
              <p>Normalized mean reciprocal rank (MRR): computed for cross-model trajectory visualization by rescaling each model’s layer-wise MRR@10 values to the 0-1 range. Raw MRR@10 values are reported alongside the normalized values in <xref rid="figure1" ref-type="fig">Figure 1</xref>.</p>
            </list-item>
          </list>
          <fig id="figure1" position="float">
            <label>Figure 1</label>
            <caption>
              <p>Layer-wise mean reciprocal rank at cutoff 10 (MRR@10) trajectories across all 13 model configurations. (A) Raw MRR@10 by relative layer depth (0=input embedding layer, 1=final layer); each line is one model configuration averaged across the 6 corpus × query-format conditions, color-coded by category. (B) The same data, minimum-maximum normalized per model to a 0-1 range to aid cross-model trajectory comparison (absolute performance differences are suppressed by the normalization). Most configurations show U-shaped or monotonically improving curves. In the per-query mixed-effects model, 3 non–retrieval-trained encoders worsen with depth (BERT-base-uncased, BioBERT, ClinicalBERT; positive rel_layer coefficients), whereas the remaining 10 configurations improve with depth, with the greatest improvement in the decoder large language models.</p>
            </caption>
            <graphic xlink:href="medinform_v14i1e99639_fig1.png" alt-version="no" mimetype="image" position="float" xlink:type="simple"/>
          </fig>
        </sec>
        <sec>
          <title>Summary Trajectory Metrics</title>
          <p>For the descriptive characterization of layer-wise trajectories, the following derived metrics are defined in the Methods section:</p>
          <list list-type="bullet">
            <list-item>
              <p>Recovery ratio: (MRR_final – MRR_trough)/(MRR_peak – MRR_trough), where MRR_peak is the highest MRR@10 among layers 0%-25% of network depth, MRR_trough is the lowest MRR@10 among layers 35%-55% of network depth, and MRR_final is the final-layer MRR@10. Values above 1.0 indicate that the final layer exceeds the early-layer peak; values below 1.0 indicate incomplete recovery. The recovery ratio is a study-defined descriptive index, not a standard metric. Per-model values are descriptive observations rather than formal statistical tests. The 1.85 threshold used for a “strong recovery” classification is data-derived and reported as a hypothesis-generating screening signal only, not a deployment criterion.</p>
            </list-item>
            <list-item>
              <p>PR expansion: PR_final/PR_trough, where the trough is the layer with the lowest within-model PR. Values above 1.0 indicate dimensionality expansion at the final layer.</p>
            </list-item>
          </list>
        </sec>
        <sec>
          <title>Document Length Analysis (Exploratory)</title>
          <p>As an exploratory analysis, documents within each corpus were partitioned into terciles based on approximate token count (whitespace-tokenized). The following token boundaries were used (33.3rd and 66.7th word-count percentiles of each corpus’s 100-document working set): MTSamples (short ≤286 words, medium 287-453, long &gt;453), PMC-Patients (short ≤331, medium 332-518, long &gt;518), and synthetic (short ≤290, medium 291-308, long &gt;308). The compressed synthetic range reflects that corpus’s structural uniformity. Per-tercile sample sizes are approximately 33 documents, and analyses at this sample size are statistically underpowered. Per-tercile retrieval metrics are reported in <xref ref-type="supplementary-material" rid="app3">Multimedia Appendix 3</xref>; the main text reports only directional summaries.</p>
        </sec>
      </sec>
      <sec>
        <title>Statistical Analysis</title>
        <p>The primary inferential analysis used a per-query, layer-wise LME model, fit separately for each of the 13 configurations. The outcome was log(rank+1), where rank is the position of the known relevant document among 100 candidates based on cosine similarity. The core model was log(rank+1) ~ C(intervention) + C(corpus) + C(query_format) + rel_layer, with a random intercept and random slope on rel_layer by corpus_query_idx (100 queries × 3 corpora = 300 groups per model). rel_layer was normalized to the range 0 to 1. Detailed coefficients, SEs, and convergence diagnostics are reported in <xref ref-type="supplementary-material" rid="app4">Multimedia Appendix 4</xref>.</p>
        <p>The primary specification was a random-slope model with a random intercept and a random slope for rel_layer by query identity. Each query was treated as a distinct group within its source corpus (300 query groups per model: 100 queries × 3 corpora). All 13 configurations converged. The estimated intraclass correlation coefficient ranged from 0.06 to 0.55 (median 0.41) across models, indicating substantial between-query variance in retrieval difficulty and supporting the per-query random-effects specification. Bootstrap CIs from 3 representative models aligned with Wald intervals in direction and approximate width. Fixed-effect <italic>P</italic> values are reported descriptively without multiple-comparison correction.</p>
        <p>No formal multiple-comparison correction was applied across layers. Layer-level <italic>P</italic> values from the per-query LMEs are interpreted descriptively within the LME inferential framework rather than as confirmatory hypothesis tests. The principal inferences in this manuscript rest on the random-slope fixed-effect coefficients and the matched-comparison contrasts, not on layer-by-layer significance counts.</p>
        <p>Cluster-bootstrap 95% CIs for the rel_layer fixed-effect coefficient were computed for 3 representative configurations spanning the model panel—BERT-base-uncased (β=+0.624, Wald 95% CI +0.537 to +0.712, bootstrap 95% CI +0.553 to +0.694), BGE-base (β=–0.165, Wald 95% CI –0.245 to –0.086, bootstrap 95% CI –0.229 to –0.104), and E5-Mistral-7B-ablation (β=–1.317, Wald 95% CI –1.392 to –1.243, bootstrap 95% CI –1.391 to –1.276)—using 50 cluster-bootstrap iterations clustered on corpus_query_idx. Bootstrap and Wald intervals agreed in sign, coefficient ordering, and approximate width across all 3 configurations; parametric SEs therefore provide a reasonable representation of uncertainty in the primary rel_layer coefficient.</p>
        <p>For the matched architectural comparisons reported in the Results section (BERT-base architecture: BGE-base vs BERT-base-uncased; Mistral-7B architecture: E5-Mistral-7B-ablation vs BioMistral-7B; biomedical encoders: BioLORD-2023 vs BioBERT), 3 pairwise pooled LME models were fit with the specification log(rank + 1) ~ rel_layer × C(model) + C(corpus) + C(query_format), with a random intercept and random slope on rel_layer by corpus_query_idx. The rel_layer × C(model) interaction term tested whether the 2 paired configurations differed in layer-depth trajectory; the reported <italic>P</italic> values are from Wald tests on the interaction coefficient. The additive per-configuration models described above and the pairwise pooled interaction models here address different inferential questions: the per-configuration slopes describe each model’s own layer-depth pattern, whereas the pooled interaction tests whether paired configurations differ in that pattern within a shared architecture.</p>
        <p>The cross-validation design used to evaluate the corpus-only ZCA intervention (post hoc interventions, below) fits the ZCA covariance from 80 training-fold documents per fold in a hidden-state space of 768-4096 dimensions, depending on the model. The resulting sample covariance is rank-deficient by 689-4017 dimensions; the whitening transform relies on the ε ridge regularizer to be well-defined outside the empirical ~80-dimensional subspace captured by the nonzero eigenvalues, where it reduces to ε-scaled near-identity behavior. The ZCA intervention in this cross-validation setup therefore operates as a hybrid of data-driven whitening within the empirical 80-dimensional subspace and ε-regularized near-identity transformation in the residual subspace. Empirically, the ε sensitivity sweep reported in <xref ref-type="supplementary-material" rid="app5">Multimedia Appendix 5</xref> confirms that the intervention’s effect on retrieval depends materially on ε within the explored range (10<sup>–7</sup> to 10<sup>–2</sup>): per-model ΔMRR@10 ranges of up to 0.75 in absolute magnitude were observed, with the sign of ΔMRR@10 changing across the sweep for several models. This is consistent with the hybrid behavior described above, where ε determines the whitening transform’s behavior in the unspanned subspace. This is a mathematical property of small-sample covariance estimation in high-dimensional spaces, not specific to the present implementation; its consequences for the interpretation of the cross-validated intervention results are discussed in the Limitations section.</p>
      </sec>
      <sec>
        <title>Post Hoc Interventions</title>
        <p>A total of 5 no-retraining interventions were evaluated across all 13 configurations: corpus-only ZCA whitening (the primary deployment-relevant analysis), transductive ZCA whitening (a query-contaminated upper-bound reference), mean centering, post hoc layer selection, and post hoc layer combination. Layer selection and layer combination are reported only as descriptive ceilings because they select layers based on test-set performance.</p>
        <p>For corpus-only ZCA, W = U(Λ + εI)^(–1/2) U^T and x’ = W(x – μ_D)/||W(x – μ_D)||, where UΛU^T is the eigendecomposition of the centered document covariance, μ_D is the document mean, and ε is a ridge regularizer. W and μ_D are fit on document embeddings only; queries are transformed afterward using the same W. The default ε was 10<sup>–5</sup>, with a sensitivity sweep over 10<sup>–7</sup> to 10<sup>–2</sup> (<xref ref-type="supplementary-material" rid="app5">Multimedia Appendix 5</xref>). The sweep showed that ε=10<sup>–5</sup> is not panel-wide optimal: tier 2 models reached their largest gains at ε in the 10<sup>–5</sup> to 10<sup>–3</sup> range (most at 10<sup>–4</sup>), whereas tier 1 models were degraded by corpus-only ZCA at every ε tested. The headline 2-tier pattern is preserved across the sweep; ε=10<sup>–5</sup> was retained as a single conservative default rather than a per-model tuned optimum.</p>
        <p>Five-fold cross-validation was performed for corpus-only ZCA by fitting W and μ_D on 80 documents and evaluating on 20 held-out documents within each corpus and query-format condition. This tests whether the transform generalizes beyond the documents used to fit it.</p>
      </sec>
      <sec>
        <title>Lexical-Overlap Audit</title>
        <p>To quantify the contribution of lexical matching to embedding-based retrieval across the 13 models, Spearman rank correlations between BM25 and embedding ranks were computed for each (model, corpus, and query format) condition. For each query, the BM25 rank assigns a position based on lexical overlap between the tokenized query and document; the embedding rank assigns a position based on cosine similarity between query and document embeddings. The Spearman correlation between these two ranks measures how well embedding-based retrieval recapitulates lexical matching: a correlation of 1.0 implies the embedding model is effectively BM25 with extra steps, while a correlation near 0 implies the embedding model uses purely nonlexical signal. Results are reported in the “Lexical Overlap Between BM25 and Embedding-Based Retrieval” subsection of the Results section and address the apparent contradiction between high BM25 baselines and the reduced-lexical-dependence claim.</p>
      </sec>
      <sec>
        <title>Chunking Sensitivity</title>
        <p>To assess sensitivity to document chunking, a sensitivity analysis was conducted on 4 representative models (BERT-base-uncased, BGE-base, MedCPT, and BioMistral-7B), each evaluated under 3 token-level truncation strategies. Documents were tokenized with the model-specific tokenizer; for the first strategy, the first N input tokens were retained, for the last strategy, the last N tokens were retained, and for the full strategy, the entire document was passed to the model with the tokenizer’s native truncation at max_length. N was set to the model’s max_length: 512 for BERT-base-uncased, BGE-base, and MedCPT; 2048 for BioMistral-7B. Because tokenizer-native truncation retained the leading max_length tokens, the full strategy produced the same leading-token input as the explicit first strategy. The last strategy therefore tested sensitivity to using document-tail rather than document-head content. Query texts remained unchanged across strategies. For each strategy, document embeddings were re-extracted, and retrieval performance was recomputed. The 4-model subset was chosen to span architecture classes: general encoder, general retriever, biomedical retriever, and biomedical LLM. Results are reported in the “Chunking Sensitivity” subsection of the Results section.</p>
      </sec>
      <sec>
        <title>Validation Corpora</title>
        <p>To evaluate generalization, the layer-level analysis was repeated on 2 independent validation corpora representing different clinical document types: 500 PMC-Patients case reports (random seed 123, no overlap with the primary 100-document sample) and 400 MTSamples medical transcriptions (the remaining documents after excluding the primary 100, drawn from the same locally cached CSV used in the companion study). Queries were generated using the same metadata-based GPT-4o approach as in the primary experiment. Validation was conducted in 2 stages. First, the 8 BERT-scale configurations from the original panel were evaluated on PMC-500 and MTSamples-400; results are reported in the “Validation on Independent Corpora” subsection of the Results section. Second, the 5 configurations not included in the first stage, BERT-base-uncased, BioMistral-7B, Phi-3-mini, E5-Mistral-7B, and E5-Mistral-7B-ablation, were evaluated on the same expanded natural-text corpora; results are reported in the “Independent Validation of the Remaining Configurations” subsection. Together, the 2 stages provide independent-corpus validation for all 13 panel configurations.</p>
      </sec>
      <sec>
        <title>Ethical Considerations</title>
        <p>This study used only publicly available, deidentified datasets and synthetic clinical notes generated by a language model. MTSamples comprises publicly posted, deidentified medical transcriptions accessed at MTSamples [<xref ref-type="bibr" rid="ref33">33</xref>] in accordance with the site’s terms of service. PMC-Patients comprises published case reports from PubMed Central [<xref ref-type="bibr" rid="ref34">34</xref>]. Synthetic clinical notes were generated by Mistral-7B-Instruct-v0.2 and contain no real patient data.</p>
        <p>No human subjects were involved in this study, and no institutional review board approval was required. No identifiable individual patient information appears in any of the corpora, the manuscript, or its supplementary materials. No compensation was provided because no participants were enrolled. As a secondary analysis of deidentified public datasets (MTSamples or PMC-Patients) and entirely model-generated synthetic notes, the study qualifies for ethics review exemption under common research ethics frameworks for non–human subject research. The original informed consent and license terms governing the source datasets cover secondary research use; no additional participant consent was applicable.</p>
      </sec>
      <sec>
        <title>Compute Environment</title>
        <p>The layer-extraction pipeline was run on Google Colab Pro+ with an NVIDIA H100 80GB GPU. Additional analyses, including matched comparisons, corpus-only ZCA cross-validation, per-query LME, lexical overlap, chunking sensitivity, and intervention sweeps, were run on Google Colab Pro+ with an NVIDIA RTX PRO 6000 Blackwell 96GB GPU. Pooling and precision choices align with the companion study configuration to ensure cross-paper reproducibility. Code is available as described in the Data Availability section.</p>
      </sec>
    </sec>
    <sec sec-type="results">
      <title>Results</title>
      <sec>
        <title>Layer-Wise Retrieval Trajectories Across 13 Configurations</title>
        <p>All 13 model configurations exhibited identifiable layer-wise structure, but the dominant pattern across the panel is more heterogeneous than that of a single U-shaped collapse model. <xref rid="figure1" ref-type="fig">Figure 1</xref> shows raw MRR@10 trajectories (panel A) and per-model minimum-maximum normalized trajectories (panel B) as a function of relative layer depth (0=embedding layer, 1=final layer), averaged across the 6 corpus × query-format conditions.</p>
        <p>Two principal trajectory patterns emerge:</p>
        <list list-type="order">
          <list-item>
            <p>Classical midlayer collapse with final-layer recovery. Observed in retrieval-trained encoder models (BGE-base, GTE-base, BioLORD-2023, MedCPT, Nomic-embed-text, and Nomic-nopfx). These models drop from early-layer MRR@10 of 0.40-0.60 to a midlayer trough of 0.04-0.25 (35%-55% relative depth), then recover sharply in the final 1-2 layers to MRR@10 of 0.80-0.88. The recovery is driven by a sharp geometric restructuring, as documented in the “Final-Layer Geometric Restructuring Is Associated With Retrieval Recovery” section.</p>
          </list-item>
          <list-item>
            <p>Monotonic improvement with layer depth is observed in most decoder-LLM configurations (E5-Mistral-7B, Phi-3-mini, BioMistral-7B, and E5-Mistral-7B-ablation). These models show retrieval performance that increases with depth, without a deep midlayer trough; final-layer values are the highest within their trajectory.</p>
          </list-item>
        </list>
        <p>A third, less common pattern, partial recovery with high final-layer anisotropy, appears in BioBERT, ClinicalBERT, and BERT-base-uncased. These 3 non–retrieval-trained encoder models show a midlayer dip but recover only modestly at the final layer (BioBERT 0.40, ClinicalBERT 0.23, and BERT-base-uncased 0.25), well below the retrieval-trained encoder family.</p>
        <p>The per-model raw and normalized values are documented in <xref ref-type="supplementary-material" rid="app6">Multimedia Appendix 6</xref>. The U-shaped trajectory is concentrated among retrieval-trained encoder models and partially evident in non–retrieval-trained encoders, but it is not characteristic of decoder LLMs. Raw MRR values and CIs are reported alongside the normalized values in <xref ref-type="supplementary-material" rid="app6">Multimedia Appendix 6</xref>.</p>
      </sec>
      <sec>
        <title>Final-Layer Trajectory Metrics</title>
        <p><xref ref-type="table" rid="table2">Table 2</xref> summarizes per-model peak, trough, and final-layer MRR@10 values, along with the recovery ratio and PR expansion factors. Complete model-level baseline and intervention results are provided in <xref ref-type="supplementary-material" rid="app7">Multimedia Appendix 7</xref>. MedCPT’s layer-0 anomaly (MRR@10=1.000, an artifact of dual-encoder shared vocabulary embeddings) is excluded from the early-peak window. The recovery ratio is structurally degenerate for MedCPT because the early-layer MRR (L1-L6 ≈ 0.025-0.033) is essentially identical to the midlayer trough, producing a near-zero denominator. We report the value with this caveat and note that MedCPT’s trajectory does not follow the U-shape geometry the metric was designed to summarize.</p>
        <p>A descriptive threshold combining recovery ratio &gt;1.85 and PR expansion &gt;1.9× is associated with final-layer MRR@10 &gt;0.80 in the retrieval-trained encoder subset (BGE-base, GTE-base, BioLORD-2023, Nomic family; all final MRR &gt;0.80, all PR expansion ≥1.90×). The matched controls and LLM-scale models confirm that this threshold is data-derived and exploratory rather than a deployment criterion.</p>
        <p>Cosine drop is negative for BioBERT (–0.030) and Phi-3-mini (–0.236), indicating that the final layer is more anisotropic than the early peak. Both fall into the extreme-anisotropy tier. The “geometric restructuring at the final layer” pattern observed in retrieval-trained encoders (positive cosine drop combined with substantial PR expansion) is therefore not a property of all transformer architectures but is specific to retrieval-objective-trained encoders.</p>
        <table-wrap position="float" id="table2">
          <label>Table 2</label>
          <caption>
            <p>Layer-level trajectory characteristics per model.<sup>a</sup></p>
          </caption>
          <table width="1000" cellpadding="5" cellspacing="0" border="1" rules="groups" frame="hsides">
            <col width="170"/>
            <col width="160"/>
            <col width="170"/>
            <col width="110"/>
            <col width="130"/>
            <col width="140"/>
            <col width="120"/>
            <thead>
              <tr valign="top">
                <td>Model</td>
                <td>Peak<sup>b</sup> (layer, MRR<sup>c</sup>)</td>
                <td>Trough<sup>d</sup> (layer, MRR)</td>
                <td>Final MRR</td>
                <td>Recovery</td>
                <td>PR expansion<sup>e</sup></td>
                <td>Cosine drop<sup>f</sup></td>
              </tr>
            </thead>
            <tbody>
              <tr valign="top">
                <td>BGE<sup>g</sup>-base</td>
                <td>L1, 0.400</td>
                <td>L4, 0.106</td>
                <td>0.872</td>
                <td>2.60</td>
                <td>2.57×</td>
                <td>+0.236</td>
              </tr>
              <tr valign="top">
                <td>GTE<sup>h</sup>-base</td>
                <td>L1, 0.408</td>
                <td>L4, 0.110</td>
                <td>0.881</td>
                <td>2.59</td>
                <td>2.63×</td>
                <td>+0.172</td>
              </tr>
              <tr valign="top">
                <td>BioLORD</td>
                <td>L0, 0.602</td>
                <td>L7, 0.414</td>
                <td>0.798</td>
                <td>2.04</td>
                <td>1.90×</td>
                <td>+0.439</td>
              </tr>
              <tr valign="top">
                <td>Nomic</td>
                <td>L1, 0.545</td>
                <td>L6, 0.225</td>
                <td>0.863</td>
                <td>2.00</td>
                <td>2.66×</td>
                <td>+0.313</td>
              </tr>
              <tr valign="top">
                <td>Nomic-nopfx</td>
                <td>L1, 0.535</td>
                <td>L6, 0.174</td>
                <td>0.863</td>
                <td>1.91</td>
                <td>2.28×</td>
                <td>+0.357</td>
              </tr>
              <tr valign="top">
                <td>MedCPT<sup>i</sup></td>
                <td>L2, 0.033</td>
                <td>L6, 0.025</td>
                <td>0.816</td>
                <td>Degenerate<sup>i</sup></td>
                <td>4.01×</td>
                <td>+0.242</td>
              </tr>
              <tr valign="top">
                <td>
                  <italic>BERT-base</italic>
                </td>
                <td>L2, 0.240</td>
                <td>L7, 0.094</td>
                <td>0.250</td>
                <td>1.07</td>
                <td>1.15×</td>
                <td>+0.009</td>
              </tr>
              <tr valign="top">
                <td>BioBERT</td>
                <td>L1, 0.350</td>
                <td>L7, 0.170</td>
                <td>0.398</td>
                <td>1.27</td>
                <td>1.35×</td>
                <td>–0.030</td>
              </tr>
              <tr valign="top">
                <td>ClinicalBERT</td>
                <td>L1, 0.167</td>
                <td>L3, 0.102</td>
                <td>0.230</td>
                <td>1.97</td>
                <td>1.38×</td>
                <td>+0.025</td>
              </tr>
              <tr valign="top">
                <td>
                  <italic>BioMistral</italic>
                </td>
                <td>L0, 0.440</td>
                <td>L11, 0.035</td>
                <td>0.424</td>
                <td>0.96</td>
                <td>1.45×</td>
                <td>+0.144</td>
              </tr>
              <tr valign="top">
                <td>Phi-3-mini</td>
                <td>L0, 0.551</td>
                <td>L12, 0.031</td>
                <td>0.522</td>
                <td>0.94</td>
                <td>2.68×</td>
                <td>–0.236</td>
              </tr>
              <tr valign="top">
                <td>E5-Mistral</td>
                <td>L0, 0.379</td>
                <td>L11, 0.044</td>
                <td>0.318</td>
                <td>0.82</td>
                <td>0.97×</td>
                <td>+0.036</td>
              </tr>
              <tr valign="top">
                <td>E5-Mistral-ablation</td>
                <td>L0, 0.430</td>
                <td>L11, 0.036</td>
                <td>0.663</td>
                <td>1.59</td>
                <td>2.24×</td>
                <td>+0.149</td>
              </tr>
            </tbody>
          </table>
          <table-wrap-foot>
            <fn id="table2fn1">
              <p><sup>a</sup>All values are averaged across 6 corpus × query-format conditions. Final MRR values match the intervention-pipeline final-layer baseline; layer-wise trajectory extraction (streams A and G) produces final-layer MRR values differing by &lt;1% from the intervention-pipeline baseline. Peak, Trough, PR expansion, and Cosine drop are computed from the layer-wise trajectory extraction. Values for new matched controls (BERT-base-uncased, BioMistral-7B) are highlighted.</p>
            </fn>
            <fn id="table2fn2">
              <p><sup>b</sup>Early peak=highest MRR@10 in the first quarter of layers (excluding layer 0 for MedCPT).</p>
            </fn>
            <fn id="table2fn3">
              <p><sup>c</sup>MRR: mean reciprocal rank.</p>
            </fn>
            <fn id="table2fn4">
              <p><sup>d</sup>Trough=lowest MRR@10 in the middle layers (35%-55% relative depth).</p>
            </fn>
            <fn id="table2fn5">
              <p><sup>e</sup>PR expansion=participation ratio at the final layer divided by participation ratio at the trough.</p>
            </fn>
            <fn id="table2fn6">
              <p><sup>f</sup>Cosine drop=average cosine at the peak layer minus average cosine at the final layer (negative values indicate increased anisotropy from peak to final).</p>
            </fn>
            <fn id="table2fn7">
              <p><sup>g</sup>BGE: Beijing Academy of Artificial Intelligence General Embedding.</p>
            </fn>
            <fn id="table2fn8">
              <p><sup>h</sup>GTE: General Text Embeddings.</p>
            </fn>
            <fn id="table2fn9">
              <p><sup>i</sup>MedCPT: layer-0 MRR@10=1.000 (excluded; reflects shared vocabulary embeddings in the dual-encoder architecture). Layers 1-6 produce near-zero MRR (range: 0.025-0.033) before a sharp recovery in the later layers; the recovery ratio, as defined in Methods, has a near-zero denominator (peak – trough ≈ 0.008) and is not meaningfully computable for this model. The PR expansion factor (4.01×) remains interpretable and is the largest in the panel.</p>
            </fn>
          </table-wrap-foot>
        </table-wrap>
      </sec>
      <sec>
        <title>Matched Architectural Comparisons Support the Training-Objective Interpretation Over Training Domain</title>
        <p>The matched comparisons supported the training objective over the training domain as the stronger explanatory axis. BERT-base-uncased and BGE-base share the BERT-base architecture but differ in retrieval fine-tuning; final-layer MRR@10 was 0.250 vs 0.872. BioMistral-7B and E5-Mistral-7B share the Mistral-7B base; the E5-Mistral ablation with mean pooling reached 0.663 vs 0.424 for BioMistral-7B. BioBERT vs BioLORD-2023 was directionally consistent (0.398 vs 0.798) but confounded by the BERT vs MPNet architecture.</p>
        <p>Formal pooled LMEs confirmed significant rel_layer × model interactions across all 3 reported contrasts (all <italic>P</italic>&lt;.001). These findings are consistent with a role for retrieval-objective training but do not establish causality because tokenizer, pretraining corpus, instruction data, pooling, and scale remain potential confounders.</p>
      </sec>
      <sec>
        <title>Final-Layer Geometric Restructuring Is Associated With Retrieval Recovery</title>
        <p>Recovery from midlayer collapse in the retrieval-trained encoder family is driven by sharp geometric restructuring in the final 1-2 layers. For BGE-base, the PR expanded substantially across the final layers (eg, from layer 10 to layer 12), while average cosine similarity declined and MRR@10 increased sharply. GTE-base, MedCPT, and Nomic-embed-text exhibited similar transition patterns. Layer-to-layer PR, cosine, and MRR values per model are reported in <xref ref-type="supplementary-material" rid="app6">Multimedia Appendix 6</xref>.</p>
        <p>In contrast, BioBERT’s PR increased modestly across layers, and its average cosine remained above 0.95 throughout. Its final-layer MRR@10 matched the value reported in <xref ref-type="table" rid="table2">Table 2</xref>. BERT-base-uncased, the matched control, exhibited similar limitations, with modest PR expansion. ClinicalBERT, with 6 transformer layers (DistilBERT architecture), lacked the depth required for a comparable recovery phase.</p>
        <p>Across the full 13-model panel, the final-layer geometric properties are summarized by category in <xref rid="figure2" ref-type="fig">Figure 2</xref>, and the average document-embedding cosine is shown in the 3-tier visualization in <xref rid="figure3" ref-type="fig">Figure 3</xref>. The matched-control evidence and geometric data converge: the retrieval-training mechanism that produces the final-layer dimensionality expansion is not produced by domain-specific pretraining alone.</p>
        <fig id="figure2" position="float">
          <label>Figure 2</label>
          <caption>
            <p>Final-layer geometric diagnostics across all 13 model configurations. (A) Participation ratio (effective embedding dimensionality), (B) average pairwise cosine similarity (a measure of representational anisotropy), and (C) anisotropy index (largest squared singular value divided by the sum of squared singular values). All quantities are computed on document embeddings and averaged across the 6 corpus × query-format conditions; points are color-coded by model category. Extreme anisotropy (average cosine &gt;0.92) appears in the non–retrieval-trained encoders and the decoder large language model Phi 3-mini (BERT-base-uncased, BioBERT, ClinicalBERT, and Phi-3-mini), whereas BioLORD-2023—whose contrastive training objective explicitly promotes isotropy—reaches the lowest average cosine in the panel (0.30).</p>
          </caption>
          <graphic xlink:href="medinform_v14i1e99639_fig2.png" alt-version="no" mimetype="image" position="float" xlink:type="simple"/>
        </fig>
        <fig id="figure3" position="float">
          <label>Figure 3</label>
          <caption>
            <p>Anisotropy tiers in final-layer document embeddings across all 13 model configurations. Final-layer document-embedding average pairwise cosine similarity, sorted in descending order, revealing 3 empirical tiers. Extreme anisotropy (&gt;0.92): Phi-3-mini (0.974), BioBERT (0.958), ClinicalBERT (0.930), and BERT-base-uncased (0.926). Moderate anisotropy (0.65-0.92): most general retrievers and the remaining large language models (LLMs). Reduced anisotropy (&lt;0.65): Nomic-embed-text-nopfx (0.638), instruction-tuned E5 Mistral-7B (0.546), and BioLORD-2023 (0.300). BioLORD-2023’s isotropy-promoting contrastive objective yields the lowest average cosine in the panel.</p>
          </caption>
          <graphic xlink:href="medinform_v14i1e99639_fig3.png" alt-version="no" mimetype="image" position="float" xlink:type="simple"/>
        </fig>
      </sec>
      <sec>
        <title>Per-Query Layer-Wise LME Analysis</title>
        <p>The per-query random-slope LME converged for all 13 configurations, each of which was fit with 300 query groups (100 queries × 3 corpora). The intraclass correlation coefficient ranged from 0.06 to 0.55 (median 0.41), indicating that between-query difficulty accounted for a substantial share of variance and supporting the use of the per-query random-effects model. Two layer-depth patterns emerged. Three non–retrieval-trained encoders worsened with depth (positive rel_layer coefficients: BERT-base-uncased +0.62, BioBERT +0.22, ClinicalBERT +0.20; all <italic>P</italic>&lt;.001), whereas the remaining 10 configurations improved with depth (negative rel_layer coefficients, all <italic>P</italic>&lt;.001), with the strongest improvement in the decoder LLMs (E5-Mistral-7B=–1.77; E5-Mistral-7B-ablation=–1.32).</p>
        <p>Additive LME coefficients for corpus-only ZCA—a layer-averaged summary—are heterogeneous across tier 1 models (negative for BGE-base, GTE-base, and Nomic-embed-text; positive for BioLORD-2023 and MedCPT) and cannot be interpreted as evidence for a uniform final-layer effect. The primary evidence for the 2-tier ZCA response is the cross-validated corpus-only ZCA analysis in <xref ref-type="table" rid="table3">Table 3</xref>, where all 6 tier 1 configurations show final-layer ΔMRR@10 losses (0.021 to 0.051) and all 7 tier 2 configurations show gains (0.066-0.304). Layer-dependent intervention effects were not interpreted from the additive LME specification. A quadratic rel_layer sensitivity check was statistically preferred across all 11 cleanly fit models, confirming that many trajectories are nonmonotonic; the linear rel_layer term is therefore interpreted as a net-direction summary. Full coefficients and diagnostics are in <xref ref-type="supplementary-material" rid="app4">Multimedia Appendix 4</xref>.</p>
        <table-wrap position="float" id="table3">
          <label>Table 3</label>
          <caption>
            <p>Five-fold cross-validation (CV): corpus-only zero-phase component analysis mean reciprocal rank (MRR) at cutoff 10 vs baseline, per model, with SE across 30 observations per model (5 folds × 6 corpus × query-format conditions). Sorted by tier and Δ.</p>
          </caption>
          <table width="1000" cellpadding="5" cellspacing="0" border="1" rules="groups" frame="hsides">
            <col width="210"/>
            <col width="200"/>
            <col width="210"/>
            <col width="120"/>
            <col width="100"/>
            <col width="80"/>
            <col width="80"/>
            <thead>
              <tr valign="top">
                <td>Model</td>
                <td>Baseline MRR</td>
                <td>ΔMRR@10, CV fold mean</td>
                <td>ΔMRR</td>
                <td>SE</td>
                <td>n</td>
                <td>Tier</td>
              </tr>
            </thead>
            <tbody>
              <tr valign="top">
                <td>BioLORD</td>
                <td>0.798</td>
                <td>0.777</td>
                <td>–0.021</td>
                <td>0.017</td>
                <td>30</td>
                <td>1</td>
              </tr>
              <tr valign="top">
                <td>MedCPT</td>
                <td>0.816</td>
                <td>0.770</td>
                <td>–0.045</td>
                <td>0.022</td>
                <td>30</td>
                <td>1</td>
              </tr>
              <tr valign="top">
                <td>BGE<sup>a</sup>-base</td>
                <td>0.872</td>
                <td>0.826</td>
                <td>–0.046</td>
                <td>0.017</td>
                <td>30</td>
                <td>1</td>
              </tr>
              <tr valign="top">
                <td>Nomic-nopfx</td>
                <td>0.867</td>
                <td>0.818</td>
                <td>–0.049</td>
                <td>0.017</td>
                <td>30</td>
                <td>1</td>
              </tr>
              <tr valign="top">
                <td>GTE<sup>b</sup>-base</td>
                <td>0.879</td>
                <td>0.830</td>
                <td>–0.050</td>
                <td>0.017</td>
                <td>30</td>
                <td>1</td>
              </tr>
              <tr valign="top">
                <td>Nomic</td>
                <td>0.862</td>
                <td>0.811</td>
                <td>–0.051</td>
                <td>0.015</td>
                <td>30</td>
                <td>1</td>
              </tr>
              <tr valign="top">
                <td>E5-Mistral</td>
                <td>0.318</td>
                <td>0.622</td>
                <td>+0.304</td>
                <td>0.023</td>
                <td>30</td>
                <td>2</td>
              </tr>
              <tr valign="top">
                <td>ClinicalBERT</td>
                <td>0.229</td>
                <td>0.436</td>
                <td>+0.207</td>
                <td>0.029</td>
                <td>30</td>
                <td>2</td>
              </tr>
              <tr valign="top">
                <td>BERT-base</td>
                <td>0.250</td>
                <td>0.449</td>
                <td>+0.200</td>
                <td>0.029</td>
                <td>30</td>
                <td>2</td>
              </tr>
              <tr valign="top">
                <td>BioBERT</td>
                <td>0.396</td>
                <td>0.592</td>
                <td>+0.196</td>
                <td>0.033</td>
                <td>30</td>
                <td>2</td>
              </tr>
              <tr valign="top">
                <td>BioMistral</td>
                <td>0.424</td>
                <td>0.585</td>
                <td>+0.161</td>
                <td>0.030</td>
                <td>30</td>
                <td>2</td>
              </tr>
              <tr valign="top">
                <td>Phi-3-mini</td>
                <td>0.523</td>
                <td>0.683</td>
                <td>+0.160</td>
                <td>0.022</td>
                <td>30</td>
                <td>2</td>
              </tr>
              <tr valign="top">
                <td>E5-Mistral-ablation</td>
                <td>0.663</td>
                <td>0.729</td>
                <td>+0.066</td>
                <td>0.032</td>
                <td>30</td>
                <td>2</td>
              </tr>
            </tbody>
          </table>
          <table-wrap-foot>
            <fn id="table3fn1">
              <p><sup>a</sup>BGE: Beijing Academy of Artificial Intelligence General Embedding.</p>
            </fn>
            <fn id="table3fn2">
              <p><sup>b</sup>GTE: General Text Embeddings.</p>
            </fn>
          </table-wrap-foot>
        </table-wrap>
      </sec>
      <sec>
        <title>Pooling and Instruction Effects in LLM-Scale Models</title>
        <p>The E5-Mistral-7B ablation provides complementary evidence on the pooling strategy. Mean pooling without an instruction prefix achieved a final-layer MRR@10 of 0.663, compared with 0.318 for vendor-specified EOS pooling with an instruction prefix; the +0.345 net improvement is consistent with the combined effects of mean pooling and removal of the instruction prefix and cannot be decomposed into pooling-only and instruction-only contributions from the available data, since both factors are confounded in the ablation. This finding motivates a deeper intervention sweep on E5-Mistral-7B-ablation.</p>
      </sec>
      <sec>
        <title>Cross-Validation of Corpus-Only ZCA Generalizes to Held-Out Documents</title>
        <p>Five-fold cross-validation was performed across all 13 models to assess the risk of test-set contamination. The ZCA transform was fit on 80 documents per fold and evaluated on the held-out 20 documents, with results ranked against the full 100-document candidate set. The mean ΔMRR@10 across folds, relative to the no-intervention baseline, is summarized in <xref rid="figure4" ref-type="fig">Figure 4</xref> and <xref ref-type="table" rid="table3">Table 3</xref>.</p>
        <p>The 2-tier pattern is clear under cross-validated evaluation. Tier 2 non–retrieval-trained models show positive ΔMRR@10 of +0.066 to +0.304 (mean +0.185 across the 7 tier 2 models). Tier 1 retrieval-trained models show negative ΔMRR@10 of –0.021 to –0.051 (mean –0.044 across the 6 tier 1 models). The cross-validation confirms that the ZCA transform fitted on 80 documents (1) generalizes to held-out documents within the same corpus, and (2) produces the expected bimodal effect: it helps the uncalibrated, modestly hurts the already-calibrated. The smallest tier 2 gain (E5-Mistral-7B-ablation=+0.066) cannot be cleanly attributed to model scale. The ablation removes the instruction prefix and switches from EOS to mean pooling simultaneously, so neither factor’s effect can be measured in isolation. Mean pooling is mathematically expected to reduce anisotropy compared with EOS pooling and may contribute to the small gain; however, the empirical anisotropy values for the 2 configurations (mean-pooled ablation 0.751; EOS-pooled instruction-tuned 0.546) represent the net difference associated with both changes simultaneously and cannot be decomposed without matched contrasts that separately vary pooling strategy and instruction formatting on the same underlying checkpoint.</p>
        <fig id="figure4" position="float">
          <label>Figure 4</label>
          <caption>
            <p>Cross-validated corpus-only zero-phase component analysis (ZCA) whitening across all 13 model configurations. Mean change in mean reciprocal rank at cutoff 10 (ΔMRR@10; corpus-only ZCA minus no-intervention baseline) per model under 5-fold cross-validation (CV). Within each fold, the ZCA transform was fit on 80 documents and evaluated on the held-out 20 documents, ranked against the full 100-document candidate set; bars show the mean across the 30 observations per model (5 folds × 6 corpus × query-format conditions), with error bars giving the standard error of that mean. The 2-tier pattern is preserved under cross-validation: tier 2 (non–retrieval-trained) models show positive deltas (+0.066 to +0.304), while tier 1 (retrieval-trained) models show small negative deltas (–0.021 to –0.051). Exact per-model values are reported in Table 3. The result confirms that the whitening transform generalizes from the fitting documents to held-out documents, helping uncalibrated encoders while modestly degrading already-calibrated retrieval models.</p>
          </caption>
          <graphic xlink:href="medinform_v14i1e99639_fig4.png" alt-version="no" mimetype="image" position="float" xlink:type="simple"/>
        </fig>
      </sec>
      <sec>
        <title>Anisotropy Across 13 Models</title>
        <p>Document-embedding geometry divided the panel into 3 tiers based on average cosine. Extreme anisotropy (&gt;0.92) was observed in Phi-3-mini (0.974), BioBERT (0.958), ClinicalBERT (0.930), and BERT-base-uncased (0.926). Moderate anisotropy (0.65-0.92) encompassed most general retrievers and LLMs. Reduced anisotropy (&lt;0.65) was observed in Nomic-embed-text-nopfx (0.638), E5-Mistral-7B with instruction (0.546), and BioLORD-2023 (0.300). Across all 13 models, the final-layer PR showed a descriptive correlation with final-layer MRR@10 (Spearman ρ=0.74; n=13). Per-condition geometry is reported in <xref ref-type="supplementary-material" rid="app6">Multimedia Appendix 6</xref>.</p>
        <p>Per-model, layer-wise Spearman correlations between PR and MRR@10, and between average pairwise cosine and MRR@10, were computed across layers for the 11 configurations included in the full per-layer geometric correlation table in <xref ref-type="supplementary-material" rid="app8">Multimedia Appendix 8</xref>. BERT-base-uncased and BioMistral-7B were extracted layer-wise for the trajectory metrics reported in <xref ref-type="table" rid="table2">Table 2</xref> and for the rel_layer LME analysis, but were not included in the <xref ref-type="supplementary-material" rid="app8">Multimedia Appendix 8</xref> correlation table. Layer-wise PR-vs-MRR Spearman ρ ranged from +0.017 (MedCPT) to +0.989 (GTE-base), with a median of +0.939; the absolute value of the corresponding cosine-vs-MRR ρ was smaller than that of PR-vs-MRR for 9 of the 11 configurations. The 2 exceptions are MedCPT, where neither correlation reaches significance because of the layer-0 dual-encoder artifact discussed earlier, and instruction-tuned E5-Mistral-7B, whose layer-wise cosine-vs-MRR ρ is positive (+0.488; <italic>P</italic>=.004), opposite in sign to the other 10 configurations and consistent with this model’s anomalous instruction-conditioned trajectory. Across the 11 configurations, the PR is the more reliable layer-wise diagnostic. The full per-model correlation table is reported in <xref ref-type="supplementary-material" rid="app8">Multimedia Appendix 8</xref>.</p>
      </sec>
      <sec>
        <title>Lexical Overlap Between BM25 and Embedding-Based Retrieval</title>
        <p>BM25-vs-embedding Spearman rank correlations were low to moderate across the panel, ranging from near zero for E5-Mistral-7B (median ρ=–0.017) to 0.366 for GTE-base. No model exceeded ρ=0.37. The high BM25 alignment baselines therefore indicate that the generated queries are aligned with their targets, while the low-to-moderate BM25-embedding correlations show that embedding retrieval is not simply lexical matching.</p>
      </sec>
      <sec>
        <title>Chunking Sensitivity</title>
        <p>For 4 representative models, BERT-base-uncased (general encoder), BGE-base (general retriever), MedCPT (biomedical retriever, dual-encoder), and BioMistral-7B (biomedical LLM), document embeddings were re-extracted using 3 chunking strategies (first-N tokens, last-N tokens, and full document, where N is each model’s maximum sequence length), and retrieval was re-evaluated. Mean MRR@10 across 6 corpus × query-format conditions are reported in <xref ref-type="table" rid="table4">Table 4</xref>.</p>
        <p>First-N and Full produce identical inputs for the three 512-ceiling models (BERT-base-uncased, BGE-base, and MedCPT) because the Full strategy relies on the tokenizer’s default front-truncation at max_length, yielding the same leading N tokens as the explicit First-N slice regardless of document length. That some documents do exceed N tokens is separately established by the Last-N vs First-N gap for these 3 models (BERT-base 0.219 vs 0.250; BGE-base 0.790 vs 0.872; MedCPT 0.782 vs 0.816); if no document exceeded 512 tokens, First-N and Last-N would return identical text for these models. For BioMistral-7B (max_length=2048), Last-N essentially matches First/Full (0.423 vs 0.424), consistent with no document exceeding 2048 tokens in the present evaluation corpora. Last-N underperforms First/Full for the three 512-ceiling encoder/retriever models by 3-8 percentage points (BERT-base 0.219 vs 0.250=–0.031; BGE-base 0.790 vs 0.872=–0.082; MedCPT 0.782 vs 0.816=–0.034) and is essentially neutral for BioMistral-7B (0.423 vs 0.424=–0.001). The pattern is consistent with clinical notes structurally front-loading key diagnostic and presenting-complaint information; encoders without retrieval-specific positional weighting perform worse when the late portion is the only input. For BioMistral-7B, the 3 strategies coincide in the present evaluation because no document required truncation within the 2048-token context window.</p>
        <p>The chunking sensitivity analysis assesses whether the principal findings are robust to alternative chunking choices. Results across the 4 representative models indicate that First-N chunking is equivalent to Full-document encoding under tokenizer-native leading truncation, and that retrieval performance is materially affected by the choice between document-head and document-tail truncation in the 512-ceiling encoder/retriever models. The principal findings are therefore not driven by the chunking choice.</p>
        <table-wrap position="float" id="table4">
          <label>Table 4</label>
          <caption>
            <p>Chunking sensitivity for 4 representative models under token-level truncation.<sup>a</sup></p>
          </caption>
          <table width="1000" cellpadding="5" cellspacing="0" border="1" rules="groups" frame="hsides">
            <col width="330"/>
            <col width="240"/>
            <col width="200"/>
            <col width="230"/>
            <thead>
              <tr valign="top">
                <td>Model</td>
                <td>First-N<sup>b</sup></td>
                <td>Full<sup>c</sup></td>
                <td>Last-N<sup>d</sup></td>
              </tr>
            </thead>
            <tbody>
              <tr valign="top">
                <td>BERT-base</td>
                <td>0.250</td>
                <td>0.250</td>
                <td>0.219</td>
              </tr>
              <tr valign="top">
                <td>BGE-base</td>
                <td>0.872</td>
                <td>0.872</td>
                <td>0.790</td>
              </tr>
              <tr valign="top">
                <td>MedCPT</td>
                <td>0.816</td>
                <td>0.816</td>
                <td>0.782</td>
              </tr>
              <tr valign="top">
                <td>BioMistral</td>
                <td>0.424</td>
                <td>0.424</td>
                <td>0.423</td>
              </tr>
            </tbody>
          </table>
          <table-wrap-foot>
            <fn id="table4fn1">
              <p><sup>a</sup>N=2048 for BioMistral-7B. Full uses tokenizer-native leading truncation at max_length; therefore, for tokenizer-default front truncation, Full and First-N produce the same leading-token input. Values are the MRR@10 across the 6 corpus × query-format conditions.</p>
            </fn>
            <fn id="table4fn2">
              <p><sup>b</sup>First-N=first N input tokens.</p>
            </fn>
            <fn id="table4fn3">
              <p><sup>c</sup>Full=full document with tokenizer-level truncation at max_length.</p>
            </fn>
            <fn id="table4fn4">
              <p><sup>d</sup>Last-N=last N input tokens. N is each model’s max_length: N=512 for BERT-base, BGE-base, and MedCPT.</p>
            </fn>
          </table-wrap-foot>
        </table-wrap>
      </sec>
      <sec>
        <title>Synthetic Corpus Distributional Audit</title>
        <p>To characterize differences between the synthetic and natural-text corpora, descriptive statistics were calculated; per-corpus document-length, vocabulary size, and type-token ratio statistics are reported in <xref ref-type="supplementary-material" rid="app9">Multimedia Appendix 9</xref>.</p>
        <p>The synthetic corpus has a mean document length that is approximately 30% shorter, a vocabulary that is 58% smaller, and a type-token ratio that is 40% lower than those of the natural-text corpora. These distributional differences explain why MedCPT and other retrieval models trained on real biomedical text underperform on synthetic data, and why corpus-only ZCA whitening is less effective on the synthetic corpus; the low-rank covariance structure of the more uniform synthetic corpus amplifies noise during whitening. This finding informs the discussion of the limitations and appropriateness of synthetic data.</p>
      </sec>
      <sec>
        <title>Post Hoc Interventions: Corpus-Only ZCA as Primary Methodology</title>
        <p>Corpus-only ZCA was the primary post hoc intervention because it applies the transform only to document embeddings. Transductive ZCA was retained as a query-contaminated reference. <xref ref-type="table" rid="table3">Table 3</xref> (cross-validated) shows the 2-tier pattern, and <xref ref-type="supplementary-material" rid="app7">Multimedia Appendix 7</xref> reports the aggregate intervention results averaged across the 6 corpus × query-format conditions: non–retrieval-trained tier 2 models generally improved under corpus-only ZCA, whereas retrieval-trained tier 1 models were degraded by additional whitening.</p>
      </sec>
      <sec>
        <title>Corpus-Dependent Collapse Patterns</title>
        <p>Midlayer collapse varied across corpora: for BioBERT, MTSamples showed a substantial decline, whereas the synthetic corpus showed minimal collapse. Natural-language queries produced deeper troughs than keyword queries but recovered to match keyword performance at the final layer. At intermediate layers, model rankings inverted relative to the final layer (Kendall τ=–0.524 between layer-9 and layer-12 rankings of the 8 BERT-scale configurations); intermediate-layer representation quality does not predict final-layer retrieval performance. Per-corpus values are reported in <xref ref-type="supplementary-material" rid="app6">Multimedia Appendix 6</xref>.</p>
      </sec>
      <sec>
        <title>Validation on Independent Corpora</title>
        <p>The first-stage validation evaluated 8 BERT-scale configurations on 500 PMC-Patients case reports and 400 MTSamples medical transcriptions. Rankings were consistent with the primary experiment (PMC-500 ρ=0.952; <italic>P</italic>&lt;.001; MTSamples-400 ρ=0.929; <italic>P</italic>=.001). Cross-source ranking between PMC case reports and medical transcriptions was high (ρ=0.976; <italic>P</italic>&lt;.001), but this primarily reflects categorical stability, as the subset separates strongly into retrieval-trained and non–retrieval-trained groups.</p>
      </sec>
      <sec>
        <title>Independent Validation of the Remaining Configurations</title>
        <p>The second-stage validation evaluated the 5 configurations not covered by the first-stage BERT-scale subset (BERT-base-uncased, BioMistral-7B, Phi-3-mini, E5-Mistral-7B, and E5-Mistral-7B-ablation)—4 LLM-scale configurations plus the BERT-base-uncased matched-comparison addition, which was not part of the original 8 BERT-scale configurations covered in the first stage—on the same expanded PMC-Patients (500) and MTSamples (400) corpora. Across the 20 evaluation conditions (5 models × 2 corpora × 2 query formats), Spearman rank correlation between primary-experiment final-layer MRR@10 and validation-corpus final-layer MRR@10 was ρ=0.839 (Kendall τ=0.632; <italic>P</italic>&lt;.001; n=20). All 20 conditions showed lower MRR@10 on the validation scale than on the primary 100-document scale, consistent with the expected effect of larger candidate pools on retrieval difficulty. Per-model mean ΔMRR@10 (averaged across the 4 corpus-by-query-format conditions per model) ranged from –0.097 for E5-Mistral-7B (smallest mean drop) to –0.247 for E5-Mistral-7B-ablation (largest), with intermediate values for BERT-base-uncased (–0.109), BioMistral-7B (–0.218), and Phi-3-mini (–0.241). The rank ordering of the 5 configurations was preserved at validation scale: E5-Mistral-7B-ablation produced the highest validation MRR@10 in all 4 corpus-by-query-format cells (matching its ordering in the primary experiment), BERT-base-uncased the lowest in all 4, and the previously reported anomaly that E5-Mistral-7B-ablation outperforms its instruction-tuned counterpart in known-item retrieval was preserved across all 4 paired validation cells. Combined with the BERT-scale validation reported above, the 2 stages together provide independent-corpus rank-stability evidence for all 13 panel configurations. Per-condition raw MRR@10 values are reported in <xref ref-type="supplementary-material" rid="app6">Multimedia Appendix 6</xref>.</p>
      </sec>
    </sec>
    <sec sec-type="discussion">
      <title>Discussion</title>
      <sec>
        <title>Principal Results</title>
        <sec>
          <title>Overview</title>
          <p>The results support 4 principal conclusions. First, anisotropy is widespread across model classes, and retrieval-objective training is associated with reduced anisotropy. Second, corpus-only ZCA explicitly reduces anisotropy and shows a 2-tier effect: it improves non–retrieval-trained models and degrades already calibrated retrieval-trained models. Third, layer-wise dynamics are heterogeneous rather than uniformly U-shaped: retrieval-trained encoders show midlayer collapse with strong final-layer recovery, non–retrieval-trained encoders recover incompletely, and decoder LLMs generally improve with depth. Fourth, BM25-overlap audits indicate that the findings are not merely lexical artifacts.</p>
          <p>The central mechanistic finding is that high final-layer retrieval performance depends on final-layer geometric restructuring rather than on the absence of midlayer collapse. In retrieval-trained encoders, midlayer troughs are followed by participation-ratio expansion and reduced cosine collapse; in non–retrieval-trained encoders, the final layer remains anisotropic, and retrieval performance remains lower. This pattern aligns with the architecture of contrastive and instruction-tuning training objectives, which compute discrimination losses at the final-layer output, making the final layer the natural locus of representation specialization. In this reading, midlayer collapse in retrieval-trained encoders is a design property of the training objective rather than a failure mode; the failure mode is the absence of subsequent final-layer geometric restructuring, which distinguishes non–retrieval-trained encoders. This framing also explains why corpus-only ZCA can approximate the geometric pattern associated with retrieval-objective training: both are associated with reshaped final-layer geometry, one through training and the other through post hoc covariance correction.</p>
        </sec>
        <sec>
          <title>Generalization to Larger and Noisier Corpora</title>
          <p>The primary 100-document known-item design is smaller and simpler than production clinical retrieval. Validation at 4-5× scale partially addresses this, but production corpora contain near-duplicates, copy-forwarded text, abbreviations, templated sections, deidentification artifacts, and institution-specific terminology. ZCA effectiveness depends on the covariance structure, so direct testing on production-scale institutional corpora remains necessary.</p>
        </sec>
      </sec>
      <sec>
        <title>Comparison With Prior Work</title>
        <p>The present analysis builds on and refines 3 threads of prior literature. First, Ethayarajh’s [<xref ref-type="bibr" rid="ref9">9</xref>] characterization of anisotropy across BERT, GPT-2, and ELMo embeddings is corroborated in the clinical retrieval setting, with 2 additions: the 13-model layer-level analysis localizes anisotropy to specific model classes (extreme in non–retrieval-trained encoders and Phi-3-mini; moderate in retrieval-trained encoders and most LLMs; reduced in BioLORD-2023 and instruction-tuned E5-Mistral-7B), and the matched-comparison evidence supports training-objective choice rather than training-domain choice as the stronger explanatory axis for the observed anisotropy differences. Second, the geometric-correction literature is extended by Su et al’s [<xref ref-type="bibr" rid="ref16">16</xref>] BERT-whitening, a transductive ZCA variant, to a deployment-relevant, corpus-only ZCA variant fitted to document embeddings without test-query information. The 2-tier pattern observed here (tier 2 models gain, tier 1 models degrade) refines the conventional framing of ZCA as a uniform improvement: the correction approximately reproduces the geometric pattern associated with retrieval-objective training. Tier membership is defined empirically by the response to whitening and, therefore, by the embedding geometry, rather than by the nominal training label. E5-Mistral-7B is the instructive exception: although it is retrieval-instruction-tuned, its decoder embeddings remain geometrically uncalibrated (final-layer average cosine 0.55, baseline MRR@10 0.32), and it shows the largest cross-validated gain in the panel (+0.304), placing it with the tier 2 models. This dissociation between the training label and geometric calibration reinforces the central claim that it is the geometry of the final-layer representation, not the training objective per se, that determines whether post hoc whitening helps. A directly parallel pattern has recently been reported in the code-search domain by Diera et al [<xref ref-type="bibr" rid="ref15">15</xref>], who applied Soft-ZCA whitening to 3 code language models (CodeBERT, CodeT5+, and Code Llama) and observed the same 2-tier effect: standard ZCA improved pretrained models substantially but degraded contrastively fine-tuned variants, with the latter requiring much smaller eigenvalue regularizers for benefit. Where retrieval-calibrated geometry is absent, ZCA tends to help; where such geometry is already present, additional whitening tends to disrupt. Third, the clinical embedding benchmarking literature, including Myers et al [<xref ref-type="bibr" rid="ref17">17</xref>] on clinical-trial retrieval, Soffer et al [<xref ref-type="bibr" rid="ref18">18</xref>] on 39 embedding models across 7 medical semantic similarity tasks reframed as retrieval problems, Khodadad et al [<xref ref-type="bibr" rid="ref35">35</xref>] on a 51-task medical embedding benchmark spanning classification, clustering, and retrieval, and the companion benchmark study [<xref ref-type="bibr" rid="ref3">3</xref>] across 3 clinical corpora and 294 conditions, is extended along 2 dimensions: layer-level analysis replaces final-layer-only evaluation, and matched-architectural controls begin to separate training-objective from training-domain contributions to performance gaps. The present manuscript is positioned as a mechanistic companion to [<xref ref-type="bibr" rid="ref3">3</xref>]: the empirical performance claim is replicated and localized rather than displaced.</p>
      </sec>
      <sec>
        <title>Limitations</title>
        <p>The study has several limitations. The primary corpora contain only 100 documents each and include a single known-relevant document per query, whereas production retrieval involves larger indices and multiple relevant documents. Queries were generated from metadata rather than written by clinicians, so they may be more label-like and less noisy than real clinical queries. The synthetic corpus differs from natural text in length, vocabulary size, and type-token ratio, limiting the generalizability of results from synthetic-only experiments.</p>
        <p>The LME uses a random-slope specification, fit separately for each of the 13 configurations, with each query treated as a distinct group within its source corpus. All models converged, showing substantial between-query variance (median intraclass correlation coefficient 0.41), computed as the variance attributable to the query random intercept divided by the total variance, per Nakagawa and Schielzeth [<xref ref-type="bibr" rid="ref36">36</xref>]. Quadratic sensitivity analyses support the substantive conclusions. Recovery ratio and PR-expansion thresholds are data-derived and should not be treated as deployment criteria. Geometry-retrieval correlations are descriptive and computed across nonindependent observations. Finally, this is a single-author study.</p>
        <p>Two further limitations and one observation merit explicit treatment. First, the known-item retrieval design assigns one relevant document per query, whereas many production clinical RAG settings involve multiple relevant documents per query, near-duplicates, and copy-forwarded notes. Multirelevant-document scoring (graded relevance, normalized discounted cumulative gain, or set-based recall) may yield different model orderings than the binary MRR@10/recall@10 used here, particularly for retrieval-trained models whose contrastive objective is optimized for single-document discrimination. Second, exploratory query-level inspection of failure cases (queries where final-layer MRR@10 was below 0.10 in the primary experiment) did not reveal a systematic pattern associated with query length, clinical specialty, or BM25-vs-embedding rank divergence; this is consistent with the BM25-vs-embedding Spearman rank correlations of –0.02 to 0.37 reported above, which indicate that embedding retrieval relies on substantial nonlexical signal that the present analysis does not further decompose. Third, the cross-source ranking stability (ρ=0.976 between PMC case reports and MTSamples medical transcriptions) primarily reflects the bimodal separation of the 8 BERT-scale subset into retrieval-trained and non–retrieval-trained groups, rather than fine-grained within-group rank stability.</p>
        <p>The corpus-only ZCA cross-validation fits the whitening transform from 80 training-fold documents per fold in a 768- to 4096-dimensional hidden-state space; the resulting sample covariance is rank-deficient by 689-4017 dimensions, and the whitening transform relies heavily on the ε ridge regularizer to be well-defined outside the empirical ~80-dimensional subspace. For the corpus covariance to be approximately full-rank (eigenvalue magnitudes dominating ε across all relevant dimensions), the number of training-fold documents must approach or exceed the hidden-state dimensionality, at least ~768 documents for BERT-scale models and ~4096 for LLM-scale. Institutional EHR corpora at the scale of 10<sup>4</sup>-10<sup>6</sup> documents are therefore expected to produce well-conditioned covariance estimates across all model classes evaluated here. The 80-document fold rank-deficiency is best understood as a methodological constraint of the diagnostic cross-validation design rather than a property of the intervention itself: at deployment scale, the covariance estimate is expected to be well-conditioned, the ε regularizer to be immaterial relative to the data-driven eigenstructure, and the ZCA intervention’s behavior to be governed by corpus geometry rather than small-sample subspace structure. Prospective evaluation at institutional scale is therefore required before any operational claim about ZCA-based correction at scale can be supported.</p>
      </sec>
      <sec>
        <title>Future Work</title>
        <p>Production-scale validation on institutional EHR corpora is the most important next step. Such validation should include (1) multirelevant-document labeling with graded relevance scoring rather than known-item retrieval; (2) clinician-authored queries reflecting real information needs rather than metadata-derived queries; (3) larger candidate indices in the 10<sup>4</sup>-10<sup>6</sup> document range; (4) near-duplicate notes, copy-forwarded text, abbreviations, templated sections, and deidentification artifacts characteristic of production EHRs; (5) temporal-drift testing across calendar periods to detect concept-drift effects on embedding geometry; (6) integration with reranking and approximate-nearest-neighbor indexing layers downstream of embedding retrieval; and (7) prospective failure analysis of retrieval misses against clinician-judged relevance. The corpus-only ZCA whitening intervention should be re-evaluated under each of these conditions before being considered for clinical deployment; the synthetic-corpus failure observed here suggests that ZCA effectiveness depends materially on the covariance structure of the deployment corpus. The matched-comparison framework introduced here should be extended to include additional clean architectural pairs as new domain-trained and retrieval-trained models become available, particularly matched LLM-scale pairs that differ only in training objective. None of the screening thresholds reported in this manuscript should be treated as deployment criteria without independent prospective validation.</p>
        <p>Matched contrasts on a single underlying model are also required to separate the pooling-attributable from the instruction-formatting-attributable components of the E5-Mistral-7B-ablation’s small post-ZCA gain. Specifically, mean-pooled and EOS-pooled representations should be compared under the same instruction-formatting protocol, and instruction-prefixed vs no-prefix representations should be compared under the same pooling strategy. These contrasts were not performed in this study and are identified as specific follow-up investigations.</p>
      </sec>
      <sec>
        <title>Conclusions</title>
        <p>Across 13 transformer configurations, retrieval performance was strongly associated with embedding geometry. Retrieval-specific training was associated with lower anisotropy, whereas corpus-only ZCA reduced anisotropy explicitly as a post hoc correction. The resulting 2-tier pattern is the central finding: non–retrieval-trained models generally improved under ZCA, whereas retrieval-trained models were degraded by additional whitening. Matched architectural contrasts support the training objective rather than the biomedical training domain as the stronger explanatory axis, although residual confounding remains.</p>
        <p>These findings provide a mechanistic basis for layer-level diagnostics and post hoc geometric correction in clinical retrieval benchmarks. They do not constitute deployment guidance. Institutional EHR validation using real queries, larger indices, multirelevant-document labels, and downstream RAG evaluation is required before clinical deployment.</p>
      </sec>
    </sec>
  </body>
  <back>
    <app-group>
      <supplementary-material id="app1">
        <label>Multimedia Appendix 1</label>
        <p>GPT-4o prompt used to generate metadata-derived queries (full prompt text, temperature, format examples).</p>
        <media xlink:href="medinform_v14i1e99639_app1.docx" xlink:title="DOCX File , 37 KB"/>
      </supplementary-material>
      <supplementary-material id="app2">
        <label>Multimedia Appendix 2</label>
        <p>Best Match 25 alignment-recovery audit for the Synthetic corpus used in the mixed-effects sensitivity check—full pair-recovery table.</p>
        <media xlink:href="medinform_v14i1e99639_app2.docx" xlink:title="DOCX File , 37 KB"/>
      </supplementary-material>
      <supplementary-material id="app3">
        <label>Multimedia Appendix 3</label>
        <p>Document-length tercile analysis with token boundaries and per-tercile retrieval metrics for the 13 models—descriptive.</p>
        <media xlink:href="medinform_v14i1e99639_app3.docx" xlink:title="DOCX File , 37 KB"/>
      </supplementary-material>
      <supplementary-material id="app4">
        <label>Multimedia Appendix 4</label>
        <p>Per-query linear mixed-effects fixed-effect coefficients and variance summary for all 13 models (each fit with 300 query groups), including fixed-effect coefficients, standard errors, z values, and <italic>P</italic> values; intraclass correlation coefficients ranged from 0.06 to 0.55 (median 0.41).</p>
        <media xlink:href="medinform_v14i1e99639_app4.docx" xlink:title="DOCX File , 35 KB"/>
      </supplementary-material>
      <supplementary-material id="app5">
        <label>Multimedia Appendix 5</label>
        <p>Epsilon sensitivity sweep results for corpus-only zero-phase component analysis across 6 ε values × 13 models × 6 conditions.</p>
        <media xlink:href="medinform_v14i1e99639_app5.docx" xlink:title="DOCX File , 29 KB"/>
      </supplementary-material>
      <supplementary-material id="app6">
        <label>Multimedia Appendix 6</label>
        <p>Per-condition raw (nonnormalized) final-layer mean reciprocal rank at cutoff 10 and recall at cutoff 10 for the 11 panel configurations included in this per-condition raw-values table; BERT-base-uncased and BioMistral-7B were extracted layer-wise for the Table 2 trajectory metrics and rel_layer linear mixed-effects analysis, but were not included in this table.</p>
        <media xlink:href="medinform_v14i1e99639_app6.docx" xlink:title="DOCX File , 134 KB"/>
      </supplementary-material>
      <supplementary-material id="app7">
        <label>Multimedia Appendix 7</label>
        <p>Per-model final-layer intervention summary: baseline, corpus-only zero-phase component analysis, and transductive zero-phase component analysis, averaged across 3 corpora × 2 query formats for all 13 models.</p>
        <media xlink:href="medinform_v14i1e99639_app7.docx" xlink:title="DOCX File , 28 KB"/>
      </supplementary-material>
      <supplementary-material id="app8">
        <label>Multimedia Appendix 8</label>
        <p>Per-model layer-wise Spearman correlations between geometric measures (participation ratio, average pairwise cosine) and mean reciprocal rank at cutoff 10 across layers, for the 11 configurations included in the full per-layer geometric correlation table. Includes PR-vs-MRR ρ, cosine-vs-MRR ρ, <italic>P</italic> values, and the per-model PR-stronger-than-cosine comparison.</p>
        <media xlink:href="medinform_v14i1e99639_app8.docx" xlink:title="DOCX File , 37 KB"/>
      </supplementary-material>
      <supplementary-material id="app9">
        <label>Multimedia Appendix 9</label>
        <p>Synthetic corpus distributional audit: per-corpus document-length, vocabulary size, and type-token ratio statistics for MTSamples, PMC-Patients, and the model-generated Synthetic corpus.</p>
        <media xlink:href="medinform_v14i1e99639_app9.docx" xlink:title="DOCX File , 37 KB"/>
      </supplementary-material>
    </app-group>
    <glossary>
      <title>Abbreviations</title>
      <def-list>
        <def-item>
          <term id="abb1">BEIR</term>
          <def>
            <p>Benchmarking Information Retrieval</p>
          </def>
        </def-item>
        <def-item>
          <term id="abb2">BGE</term>
          <def>
            <p>Beijing Academy of Artificial Intelligence General Embedding</p>
          </def>
        </def-item>
        <def-item>
          <term id="abb3">BM25</term>
          <def>
            <p>Best Match 25</p>
          </def>
        </def-item>
        <def-item>
          <term id="abb4">EHR</term>
          <def>
            <p>electronic health record</p>
          </def>
        </def-item>
        <def-item>
          <term id="abb5">EOS</term>
          <def>
            <p>end of sequence</p>
          </def>
        </def-item>
        <def-item>
          <term id="abb6">GTE</term>
          <def>
            <p>General Text Embeddings</p>
          </def>
        </def-item>
        <def-item>
          <term id="abb7">LLM</term>
          <def>
            <p>large language model</p>
          </def>
        </def-item>
        <def-item>
          <term id="abb8">LME</term>
          <def>
            <p>linear mixed-effects</p>
          </def>
        </def-item>
        <def-item>
          <term id="abb9">MRR</term>
          <def>
            <p>mean reciprocal rank</p>
          </def>
        </def-item>
        <def-item>
          <term id="abb10">MRR@10</term>
          <def>
            <p>mean reciprocal rank at cutoff 10</p>
          </def>
        </def-item>
        <def-item>
          <term id="abb11">PR</term>
          <def>
            <p>participation ratio</p>
          </def>
        </def-item>
        <def-item>
          <term id="abb12">RAG</term>
          <def>
            <p>retrieval-augmented generation</p>
          </def>
        </def-item>
        <def-item>
          <term id="abb13">Recall@10</term>
          <def>
            <p>recall at cutoff 10</p>
          </def>
        </def-item>
        <def-item>
          <term id="abb14">ZCA</term>
          <def>
            <p>zero-phase component analysis</p>
          </def>
        </def-item>
      </def-list>
    </glossary>
    <ack>
      <p>The author thanks the JMIR Medical Informatics editorial team and the peer reviewers for constructive comments that improved the manuscript.</p>
      <p>Mistral-7B-Instruct-v0.2 (mistralai, accessed via Hugging Face) was used in the companion study [<xref ref-type="bibr" rid="ref3">3</xref>] to generate the synthetic clinical notes reused here. GPT-4o (gpt-4o-2024-05-13, OpenAI) was used to generate metadata-derived queries (temperature 0.3) following the procedure described in the companion study; the full prompt is provided in <xref ref-type="supplementary-material" rid="app1">Multimedia Appendix 1</xref>. Claude (Anthropic) served as a coding assistant for compute scripts and statistical analyses and as a writing assistant for initial drafting and revision of the manuscript text. Grammarly was used for spelling and grammar. All code was reviewed, validated, and executed by the author. The manuscript was critically reviewed, revised, and finalized by the author, who takes full responsibility for the scientific content, interpretation, and conclusions.</p>
    </ack>
    <notes>
      <sec>
        <title>Funding</title>
        <p>This research received no specific grant from any funding agency in the public, commercial, or nonprofit sectors. Computational resources were provided via Google Colab Pro+.</p>
      </sec>
    </notes>
    <notes>
      <sec>
        <title>Data Availability</title>
        <p>Code, layer-level results, intervention data, validation results, analysis scripts, and figure-generation code are available in the GitHub repository for this study, with archived snapshots at Zenodo [<xref ref-type="bibr" rid="ref37">37</xref>] under a CC BY 4.0 license. The compute scripts and outputs are included in the same repository. Query sets, synthetic clinical notes, and the MTSamples cache used here are available in the companion study repository [<xref ref-type="bibr" rid="ref38">38</xref>]. The ZCA whitening implementation evaluated here is available as the clinical-embedding-fix software library, deposited on Zenodo [<xref ref-type="bibr" rid="ref39">39</xref>]. PMC-Patients data can be reproduced from the zhengyun21/PMC-Patients dataset using the seeds specified in the analysis scripts.</p>
      </sec>
    </notes>
    <fn-group>
      <fn fn-type="con">
        <p>YM contributed to conceptualization, methodology, software, validation, formal analysis, investigation, data curation, writing—original draft, writing—review and editing, and visualization. The author read and approved the final manuscript.</p>
      </fn>
      <fn fn-type="conflict">
        <p>The author serves as the Chief Medical Officer of AlgiPharma AS. AlgiPharma has no products or commercial interests related to embedding models, retrieval-augmented generation, or clinical informatics systems. No other competing interests are disclosed.</p>
      </fn>
    </fn-group>
    <ref-list>
      <ref id="ref1">
        <label>1</label>
        <nlm-citation citation-type="journal">
          <person-group person-group-type="author">
            <name name-style="western">
              <surname>Lewis</surname>
              <given-names>P</given-names>
            </name>
            <name name-style="western">
              <surname>Perez</surname>
              <given-names>E</given-names>
            </name>
            <name name-style="western">
              <surname>Piktus</surname>
              <given-names>A</given-names>
            </name>
            <name name-style="western">
              <surname>Petroni</surname>
              <given-names>F</given-names>
            </name>
            <name name-style="western">
              <surname>Karpukhin</surname>
              <given-names>V</given-names>
            </name>
            <name name-style="western">
              <surname>Goyal</surname>
              <given-names>N</given-names>
            </name>
            <name name-style="western">
              <surname>Küttler</surname>
              <given-names>H</given-names>
            </name>
          </person-group>
          <article-title>Retrieval-augmented generation for knowledge-intensive NLP tasks</article-title>
          <source>arXiv</source>
          <comment>Preprint posted online on May 22, 2020</comment>
          <pub-id pub-id-type="doi">10.48550/arXiv.2005.11401</pub-id>
        </nlm-citation>
      </ref>
      <ref id="ref2">
        <label>2</label>
        <nlm-citation citation-type="journal">
          <person-group person-group-type="author">
            <name name-style="western">
              <surname>Gao</surname>
              <given-names>Y</given-names>
            </name>
            <name name-style="western">
              <surname>Xiong</surname>
              <given-names>Y</given-names>
            </name>
            <name name-style="western">
              <surname>Gao</surname>
              <given-names>X</given-names>
            </name>
            <name name-style="western">
              <surname>Jia</surname>
              <given-names>K</given-names>
            </name>
            <name name-style="western">
              <surname>Pan</surname>
              <given-names>J</given-names>
            </name>
            <name name-style="western">
              <surname>Bi</surname>
              <given-names>Y</given-names>
            </name>
            <name name-style="western">
              <surname>Dai</surname>
              <given-names>Y</given-names>
            </name>
          </person-group>
          <article-title>Retrieval-augmented generation for large language models: a survey</article-title>
          <source>arXiv</source>
          <comment>Preprint posted online on December 18, 2023</comment>
          <pub-id pub-id-type="doi">10.48550/arXiv.2312.10997</pub-id>
        </nlm-citation>
      </ref>
      <ref id="ref3">
        <label>3</label>
        <nlm-citation citation-type="journal">
          <person-group person-group-type="author">
            <name name-style="western">
              <surname>Mikkelsen</surname>
              <given-names>Y</given-names>
            </name>
          </person-group>
          <article-title>Clinical context variables collectively rival model choice in embedding-based retrieval: multi-corpus benchmark study</article-title>
          <source>JMIR Med Inform</source>
          <year>2026</year>
          <volume>14</volume>
          <fpage>e94241</fpage>
          <comment>
            <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://medinform.jmir.org/2026//e94241/"/>
          </comment>
          <pub-id pub-id-type="doi">10.2196/94241</pub-id>
          <pub-id pub-id-type="medline">42097608</pub-id>
          <pub-id pub-id-type="pii">v14i1e94241</pub-id>
          <pub-id pub-id-type="pmcid">PMC13195371</pub-id>
        </nlm-citation>
      </ref>
      <ref id="ref4">
        <label>4</label>
        <nlm-citation citation-type="journal">
          <person-group person-group-type="author">
            <name name-style="western">
              <surname>Lee</surname>
              <given-names>J</given-names>
            </name>
            <name name-style="western">
              <surname>Yoon</surname>
              <given-names>W</given-names>
            </name>
            <name name-style="western">
              <surname>Kim</surname>
              <given-names>S</given-names>
            </name>
            <name name-style="western">
              <surname>Kim</surname>
              <given-names>D</given-names>
            </name>
            <name name-style="western">
              <surname>Kim</surname>
              <given-names>S</given-names>
            </name>
            <name name-style="western">
              <surname>So</surname>
              <given-names>CH</given-names>
            </name>
            <name name-style="western">
              <surname>Kang</surname>
              <given-names>J</given-names>
            </name>
          </person-group>
          <article-title>BioBERT: a pre-trained biomedical language representation model for biomedical text mining</article-title>
          <source>Bioinformatics</source>
          <year>2020</year>
          <volume>36</volume>
          <issue>4</issue>
          <fpage>1234</fpage>
          <lpage>1240</lpage>
          <comment>
            <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://europepmc.org/abstract/MED/31501885"/>
          </comment>
          <pub-id pub-id-type="doi">10.1093/bioinformatics/btz682</pub-id>
          <pub-id pub-id-type="medline">31501885</pub-id>
          <pub-id pub-id-type="pii">5566506</pub-id>
          <pub-id pub-id-type="pmcid">PMC7703786</pub-id>
        </nlm-citation>
      </ref>
      <ref id="ref5">
        <label>5</label>
        <nlm-citation citation-type="confproc">
          <person-group person-group-type="author">
            <name name-style="western">
              <surname>Alsentzer</surname>
              <given-names>E</given-names>
            </name>
            <name name-style="western">
              <surname>Murphy</surname>
              <given-names>J</given-names>
            </name>
            <name name-style="western">
              <surname>Boag</surname>
              <given-names>W</given-names>
            </name>
          </person-group>
          <article-title>Publicly available clinical BERT embeddings</article-title>
          <year>2019</year>
          <conf-name>Proceedings of the 2nd Clinical Natural Language Processing Workshop</conf-name>
          <conf-date>June 7, 2019</conf-date>
          <conf-loc>Minneapolis, MN</conf-loc>
          <fpage>72</fpage>
          <lpage>78</lpage>
          <pub-id pub-id-type="doi">10.18653/v1/w19-1909</pub-id>
        </nlm-citation>
      </ref>
      <ref id="ref6">
        <label>6</label>
        <nlm-citation citation-type="confproc">
          <person-group person-group-type="author">
            <name name-style="western">
              <surname>Xiao</surname>
              <given-names>S</given-names>
            </name>
            <name name-style="western">
              <surname>Liu</surname>
              <given-names>Z</given-names>
            </name>
            <name name-style="western">
              <surname>Zhang</surname>
              <given-names>P</given-names>
            </name>
          </person-group>
          <article-title>C-Pack: packaged resources to advance general Chinese embedding</article-title>
          <year>2024</year>
          <conf-name>SIGIR '24: Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval</conf-name>
          <conf-date>July 14-18, 2024</conf-date>
          <conf-loc>Washington, DC</conf-loc>
          <fpage>641</fpage>
          <lpage>649</lpage>
          <comment>
            <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://doi.org/10.48550/arXiv.2309.07597"/>
          </comment>
        </nlm-citation>
      </ref>
      <ref id="ref7">
        <label>7</label>
        <nlm-citation citation-type="journal">
          <person-group person-group-type="author">
            <name name-style="western">
              <surname>Li</surname>
              <given-names>Z</given-names>
            </name>
            <name name-style="western">
              <surname>Zhang</surname>
              <given-names>X</given-names>
            </name>
            <name name-style="western">
              <surname>Zhang</surname>
              <given-names>Y</given-names>
            </name>
            <name name-style="western">
              <surname>Long</surname>
              <given-names>D</given-names>
            </name>
            <name name-style="western">
              <surname>Xie</surname>
              <given-names>P</given-names>
            </name>
            <name name-style="western">
              <surname>Zhang</surname>
              <given-names>M</given-names>
            </name>
          </person-group>
          <article-title>Towards general text embeddings with multi-stage contrastive learning</article-title>
          <source>arXiv</source>
          <comment>Preprint posted online on August 7, 2023</comment>
          <pub-id pub-id-type="doi">10.48550/arXiv.2308.03281</pub-id>
        </nlm-citation>
      </ref>
      <ref id="ref8">
        <label>8</label>
        <nlm-citation citation-type="journal">
          <person-group person-group-type="author">
            <name name-style="western">
              <surname>Nussbaum</surname>
              <given-names>Z</given-names>
            </name>
            <name name-style="western">
              <surname>Morris</surname>
              <given-names>J</given-names>
            </name>
            <name name-style="western">
              <surname>Duderstadt</surname>
              <given-names>B</given-names>
            </name>
            <name name-style="western">
              <surname>Mulyar</surname>
              <given-names>A</given-names>
            </name>
          </person-group>
          <article-title>Nomic embed: training a reproducible long context text embedder</article-title>
          <source>arXiv</source>
          <comment>Preprint posted online on February 2, 2024</comment>
          <pub-id pub-id-type="doi">10.48550/arXiv.2402.01613</pub-id>
        </nlm-citation>
      </ref>
      <ref id="ref9">
        <label>9</label>
        <nlm-citation citation-type="confproc">
          <person-group person-group-type="author">
            <name name-style="western">
              <surname>Ethayarajh</surname>
              <given-names>K</given-names>
            </name>
          </person-group>
          <article-title>How contextual are contextualized word representations? Comparing the geometry of BERT, ELMo, and GPT-2 embeddings</article-title>
          <year>2019</year>
          <conf-name>Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)</conf-name>
          <conf-date>July 16, 2019</conf-date>
          <conf-loc>Hong Kong, China</conf-loc>
          <fpage>55</fpage>
          <lpage>65</lpage>
          <pub-id pub-id-type="doi">10.18653/v1/d19-1006</pub-id>
        </nlm-citation>
      </ref>
      <ref id="ref10">
        <label>10</label>
        <nlm-citation citation-type="confproc">
          <person-group person-group-type="author">
            <name name-style="western">
              <surname>Razzhigaev</surname>
              <given-names>A</given-names>
            </name>
            <name name-style="western">
              <surname>Mikhalchuk</surname>
              <given-names>M</given-names>
            </name>
            <name name-style="western">
              <surname>Goncharova</surname>
              <given-names>E</given-names>
            </name>
            <name name-style="western">
              <surname>Oseledets</surname>
              <given-names>I</given-names>
            </name>
            <name name-style="western">
              <surname>Dimitrov</surname>
              <given-names>D</given-names>
            </name>
            <name name-style="western">
              <surname>Kuznetsov</surname>
              <given-names>A</given-names>
            </name>
          </person-group>
          <article-title>The shape of learning: anisotropy and intrinsic dimensions in transformer-based models</article-title>
          <year>2024</year>
          <conf-name>Findings of the Association for Computational Linguistics: EACL 2024</conf-name>
          <conf-date>March 17-22, 2024</conf-date>
          <conf-loc>St. Julian’s, Malta</conf-loc>
          <fpage>868</fpage>
          <lpage>874</lpage>
          <pub-id pub-id-type="doi">10.18653/v1/2024.findings-eacl.58</pub-id>
        </nlm-citation>
      </ref>
      <ref id="ref11">
        <label>11</label>
        <nlm-citation citation-type="confproc">
          <person-group person-group-type="author">
            <name name-style="western">
              <surname>Machina</surname>
              <given-names>A</given-names>
            </name>
            <name name-style="western">
              <surname>Mercer</surname>
              <given-names>R</given-names>
            </name>
          </person-group>
          <article-title>Anisotropy is not inherent to transformers</article-title>
          <year>2024</year>
          <conf-name>Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)</conf-name>
          <conf-date>July 16-21, 2024</conf-date>
          <conf-loc>Mexico City, Mexico</conf-loc>
          <fpage>4892</fpage>
          <lpage>4907</lpage>
          <pub-id pub-id-type="doi">10.18653/v1/2024.naacl-long.274</pub-id>
        </nlm-citation>
      </ref>
      <ref id="ref12">
        <label>12</label>
        <nlm-citation citation-type="confproc">
          <person-group person-group-type="author">
            <name name-style="western">
              <surname>Godey</surname>
              <given-names>N</given-names>
            </name>
            <name name-style="western">
              <surname>Clergerie</surname>
              <given-names>È</given-names>
            </name>
            <name name-style="western">
              <surname>Sagot</surname>
              <given-names>B</given-names>
            </name>
          </person-group>
          <article-title>Anisotropy is inherent to self-attention in transformers</article-title>
          <year>2024</year>
          <conf-name>Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers)</conf-name>
          <conf-date>March 17-22, 2024</conf-date>
          <conf-loc>St. Julian’s, Malta</conf-loc>
          <fpage>35</fpage>
          <lpage>48</lpage>
          <pub-id pub-id-type="doi">10.18653/v1/2024.eacl-long.3</pub-id>
        </nlm-citation>
      </ref>
      <ref id="ref13">
        <label>13</label>
        <nlm-citation citation-type="journal">
          <person-group person-group-type="author">
            <name name-style="western">
              <surname>Remy</surname>
              <given-names>F</given-names>
            </name>
            <name name-style="western">
              <surname>Demuynck</surname>
              <given-names>K</given-names>
            </name>
            <name name-style="western">
              <surname>Demeester</surname>
              <given-names>T</given-names>
            </name>
          </person-group>
          <article-title>BioLORD-2023: semantic textual representations fusing large language models and clinical knowledge graph insights</article-title>
          <source>J Am Med Inform Assoc</source>
          <year>2024</year>
          <volume>31</volume>
          <issue>9</issue>
          <fpage>1844</fpage>
          <lpage>1855</lpage>
          <comment>
            <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://academic.oup.com/jamia/article-lookup/doi/10.1093/jamia/ocae029"/>
          </comment>
          <pub-id pub-id-type="doi">10.1093/jamia/ocae029</pub-id>
          <pub-id pub-id-type="medline">38412333</pub-id>
          <pub-id pub-id-type="pii">7614965</pub-id>
          <pub-id pub-id-type="pmcid">PMC11339519</pub-id>
        </nlm-citation>
      </ref>
      <ref id="ref14">
        <label>14</label>
        <nlm-citation citation-type="journal">
          <person-group person-group-type="author">
            <name name-style="western">
              <surname>Bell</surname>
              <given-names>AJ</given-names>
            </name>
            <name name-style="western">
              <surname>Sejnowski</surname>
              <given-names>TJ</given-names>
            </name>
          </person-group>
          <article-title>The "independent components" of natural scenes are edge filters</article-title>
          <source>Vision Res</source>
          <year>1997</year>
          <volume>37</volume>
          <issue>23</issue>
          <fpage>3327</fpage>
          <lpage>3338</lpage>
          <comment>
            <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://linkinghub.elsevier.com/retrieve/pii/S0042-6989(97)00121-1"/>
          </comment>
          <pub-id pub-id-type="doi">10.1016/s0042-6989(97)00121-1</pub-id>
          <pub-id pub-id-type="medline">9425547</pub-id>
          <pub-id pub-id-type="pii">S0042-6989(97)00121-1</pub-id>
          <pub-id pub-id-type="pmcid">PMC2882863</pub-id>
        </nlm-citation>
      </ref>
      <ref id="ref15">
        <label>15</label>
        <nlm-citation citation-type="journal">
          <person-group person-group-type="author">
            <name name-style="western">
              <surname>Diera</surname>
              <given-names>A</given-names>
            </name>
            <name name-style="western">
              <surname>Galke</surname>
              <given-names>L</given-names>
            </name>
            <name name-style="western">
              <surname>Scherp</surname>
              <given-names>A</given-names>
            </name>
          </person-group>
          <article-title>Isotropy matters: soft-ZCA whitening of embeddings for semantic code search</article-title>
          <source>arXiv</source>
          <comment>Preprint posted online on November 26, 2024</comment>
          <pub-id pub-id-type="doi">10.48550/arXiv.2411.17538</pub-id>
        </nlm-citation>
      </ref>
      <ref id="ref16">
        <label>16</label>
        <nlm-citation citation-type="journal">
          <person-group person-group-type="author">
            <name name-style="western">
              <surname>Su</surname>
              <given-names>J</given-names>
            </name>
            <name name-style="western">
              <surname>Cao</surname>
              <given-names>J</given-names>
            </name>
            <name name-style="western">
              <surname>Liu</surname>
              <given-names>W</given-names>
            </name>
            <name name-style="western">
              <surname>Ou</surname>
              <given-names>Y</given-names>
            </name>
          </person-group>
          <article-title>Whitening sentence representations for better semantics and faster retrieval</article-title>
          <source>arXiv</source>
          <comment>Preprint posted online on March 29, 2021</comment>
          <pub-id pub-id-type="doi">10.48550/arXiv.2103.15316</pub-id>
        </nlm-citation>
      </ref>
      <ref id="ref17">
        <label>17</label>
        <nlm-citation citation-type="journal">
          <person-group person-group-type="author">
            <name name-style="western">
              <surname>Myers</surname>
              <given-names>S</given-names>
            </name>
            <name name-style="western">
              <surname>Miller</surname>
              <given-names>TA</given-names>
            </name>
            <name name-style="western">
              <surname>Gao</surname>
              <given-names>Y</given-names>
            </name>
            <name name-style="western">
              <surname>Churpek</surname>
              <given-names>MM</given-names>
            </name>
            <name name-style="western">
              <surname>Mayampurath</surname>
              <given-names>A</given-names>
            </name>
            <name name-style="western">
              <surname>Dligach</surname>
              <given-names>D</given-names>
            </name>
            <name name-style="western">
              <surname>Afshar</surname>
              <given-names>M</given-names>
            </name>
          </person-group>
          <article-title>Lessons learned on information retrieval in electronic health records: a comparison of embedding models and pooling strategies</article-title>
          <source>J Am Med Inform Assoc</source>
          <year>2025</year>
          <volume>32</volume>
          <issue>2</issue>
          <fpage>357</fpage>
          <lpage>364</lpage>
          <pub-id pub-id-type="doi">10.1093/jamia/ocae308</pub-id>
          <pub-id pub-id-type="medline">39703187</pub-id>
          <pub-id pub-id-type="pii">7929311</pub-id>
          <pub-id pub-id-type="pmcid">PMC11756698</pub-id>
        </nlm-citation>
      </ref>
      <ref id="ref18">
        <label>18</label>
        <nlm-citation citation-type="journal">
          <person-group person-group-type="author">
            <name name-style="western">
              <surname>Soffer</surname>
              <given-names>S</given-names>
            </name>
            <name name-style="western">
              <surname>Omar</surname>
              <given-names>M</given-names>
            </name>
            <name name-style="western">
              <surname>Gendler</surname>
              <given-names>M</given-names>
            </name>
            <name name-style="western">
              <surname>Glicksberg</surname>
              <given-names>BS</given-names>
            </name>
            <name name-style="western">
              <surname>Kovatch</surname>
              <given-names>P</given-names>
            </name>
            <name name-style="western">
              <surname>Efros</surname>
              <given-names>O</given-names>
            </name>
            <name name-style="western">
              <surname>Freeman</surname>
              <given-names>R</given-names>
            </name>
            <name name-style="western">
              <surname>Charney</surname>
              <given-names>AW</given-names>
            </name>
            <name name-style="western">
              <surname>Nadkarni</surname>
              <given-names>GN</given-names>
            </name>
            <name name-style="western">
              <surname>Klang</surname>
              <given-names>E</given-names>
            </name>
          </person-group>
          <article-title>A scalable framework for benchmark embedding models in semantic health-care tasks</article-title>
          <source>J Am Med Inform Assoc</source>
          <year>2025</year>
          <volume>32</volume>
          <issue>12</issue>
          <fpage>1877</fpage>
          <lpage>1887</lpage>
          <pub-id pub-id-type="doi">10.1093/jamia/ocaf149</pub-id>
          <pub-id pub-id-type="medline">40977370</pub-id>
          <pub-id pub-id-type="pii">8261126</pub-id>
          <pub-id pub-id-type="pmcid">PMC12646376</pub-id>
        </nlm-citation>
      </ref>
      <ref id="ref19">
        <label>19</label>
        <nlm-citation citation-type="confproc">
          <person-group person-group-type="author">
            <name name-style="western">
              <surname>Hosseini</surname>
              <given-names>M</given-names>
            </name>
            <name name-style="western">
              <surname>Munia</surname>
              <given-names>M</given-names>
            </name>
            <name name-style="western">
              <surname>Khan</surname>
              <given-names>L</given-names>
            </name>
          </person-group>
          <article-title>BERT has more to offer: BERT layers combination yields better sentence embedding</article-title>
          <year>2023</year>
          <conf-name>Findings of the Association for Computational Linguistics: EMNLP 2023</conf-name>
          <conf-date>July 16, 2026</conf-date>
          <conf-loc>Singapore</conf-loc>
          <fpage>15419</fpage>
          <lpage>15431</lpage>
          <pub-id pub-id-type="doi">10.18653/v1/2023.findings-emnlp.1030</pub-id>
        </nlm-citation>
      </ref>
      <ref id="ref20">
        <label>20</label>
        <nlm-citation citation-type="confproc">
          <person-group person-group-type="author">
            <name name-style="western">
              <surname>Muennighoff</surname>
              <given-names>N</given-names>
            </name>
            <name name-style="western">
              <surname>Tazi</surname>
              <given-names>N</given-names>
            </name>
            <name name-style="western">
              <surname>Magne</surname>
              <given-names>L</given-names>
            </name>
            <name name-style="western">
              <surname>Reimers</surname>
              <given-names>N</given-names>
            </name>
          </person-group>
          <article-title>MTEB: massive text embedding benchmark</article-title>
          <year>2023</year>
          <conf-name>Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics</conf-name>
          <conf-date>May 2-6, 2023</conf-date>
          <conf-loc>Dubrovnik, Croatia</conf-loc>
          <fpage>2014</fpage>
          <lpage>2037</lpage>
          <pub-id pub-id-type="doi">10.18653/v1/2023.eacl-main.148</pub-id>
        </nlm-citation>
      </ref>
      <ref id="ref21">
        <label>21</label>
        <nlm-citation citation-type="journal">
          <person-group person-group-type="author">
            <name name-style="western">
              <surname>Thakur</surname>
              <given-names>N</given-names>
            </name>
            <name name-style="western">
              <surname>Reimers</surname>
              <given-names>N</given-names>
            </name>
            <name name-style="western">
              <surname>Rücklé</surname>
              <given-names>A</given-names>
            </name>
            <name name-style="western">
              <surname>Srivastava</surname>
              <given-names>A</given-names>
            </name>
            <name name-style="western">
              <surname>Gurevych</surname>
              <given-names>I</given-names>
            </name>
          </person-group>
          <article-title>BEIR: a heterogenous benchmark for zero-shot evaluation of information retrieval models</article-title>
          <source>arXiv</source>
          <comment>Preprint posted online on April 17, 2021</comment>
          <pub-id pub-id-type="doi">10.48550/arXiv.2104.08663</pub-id>
        </nlm-citation>
      </ref>
      <ref id="ref22">
        <label>22</label>
        <nlm-citation citation-type="confproc">
          <person-group person-group-type="author">
            <name name-style="western">
              <surname>Karpukhin</surname>
              <given-names>V</given-names>
            </name>
            <name name-style="western">
              <surname>Oguz</surname>
              <given-names>B</given-names>
            </name>
            <name name-style="western">
              <surname>Min</surname>
              <given-names>S</given-names>
            </name>
            <name name-style="western">
              <surname>Lewis</surname>
              <given-names>P</given-names>
            </name>
            <name name-style="western">
              <surname>Wu</surname>
              <given-names>L</given-names>
            </name>
            <name name-style="western">
              <surname>Edunov</surname>
              <given-names>S</given-names>
            </name>
          </person-group>
          <article-title>Dense passage retrieval for open-domain question answering</article-title>
          <year>2020</year>
          <conf-name>Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)</conf-name>
          <conf-date>November 16-20, 2020</conf-date>
          <conf-loc>Online</conf-loc>
          <fpage>6769</fpage>
          <lpage>6781</lpage>
          <pub-id pub-id-type="doi">10.18653/v1/2020.emnlp-main.550</pub-id>
        </nlm-citation>
      </ref>
      <ref id="ref23">
        <label>23</label>
        <nlm-citation citation-type="confproc">
          <person-group person-group-type="author">
            <name name-style="western">
              <surname>Sciavolino</surname>
              <given-names>C</given-names>
            </name>
            <name name-style="western">
              <surname>Zhong</surname>
              <given-names>Z</given-names>
            </name>
            <name name-style="western">
              <surname>Lee</surname>
              <given-names>J</given-names>
            </name>
            <name name-style="western">
              <surname>Chen</surname>
              <given-names>D</given-names>
            </name>
          </person-group>
          <article-title>Simple entity-centric questions challenge dense retrievers</article-title>
          <year>2021</year>
          <conf-name>Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing</conf-name>
          <conf-date>November 7-11, 2021</conf-date>
          <conf-loc>Punta Cana, Dominican Republic</conf-loc>
          <fpage>6138</fpage>
          <lpage>6148</lpage>
          <pub-id pub-id-type="doi">10.18653/v1/2021.emnlp-main.496</pub-id>
        </nlm-citation>
      </ref>
      <ref id="ref24">
        <label>24</label>
        <nlm-citation citation-type="confproc">
          <person-group person-group-type="author">
            <name name-style="western">
              <surname>Khattab</surname>
              <given-names>O</given-names>
            </name>
            <name name-style="western">
              <surname>Zaharia</surname>
              <given-names>M</given-names>
            </name>
          </person-group>
          <article-title>ColBERT: efficient and effective passage search via contextualized late interaction over BERT</article-title>
          <year>2020</year>
          <conf-name>SIGIR '20: Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval</conf-name>
          <conf-date>July 25-30, 2020</conf-date>
          <conf-loc>Virtual Event, China</conf-loc>
          <fpage>39</fpage>
          <lpage>48</lpage>
          <pub-id pub-id-type="doi">10.1145/3397271.3401075</pub-id>
        </nlm-citation>
      </ref>
      <ref id="ref25">
        <label>25</label>
        <nlm-citation citation-type="journal">
          <person-group person-group-type="author">
            <name name-style="western">
              <surname>Nogueira</surname>
              <given-names>R</given-names>
            </name>
            <name name-style="western">
              <surname>Cho</surname>
              <given-names>K</given-names>
            </name>
          </person-group>
          <article-title>Passage re-ranking with BERT</article-title>
          <source>arXiv</source>
          <comment>Preprint posted online on January 13, 2019</comment>
          <pub-id pub-id-type="doi">10.48550/arXiv.1901.04085</pub-id>
        </nlm-citation>
      </ref>
      <ref id="ref26">
        <label>26</label>
        <nlm-citation citation-type="confproc">
          <person-group person-group-type="author">
            <name name-style="western">
              <surname>Wang</surname>
              <given-names>K</given-names>
            </name>
            <name name-style="western">
              <surname>Reimers</surname>
              <given-names>N</given-names>
            </name>
            <name name-style="western">
              <surname>Gurevych</surname>
              <given-names>I</given-names>
            </name>
          </person-group>
          <article-title>TSDAE: using transformer-based sequential denoising auto-encoder for unsupervised sentence embedding learning</article-title>
          <year>2021</year>
          <conf-name>Findings of the Association for Computational Linguistics: EMNLP 2021</conf-name>
          <conf-date>November 7-11, 2021</conf-date>
          <conf-loc>Punta Cana, Dominican Republic</conf-loc>
          <fpage>671</fpage>
          <lpage>688</lpage>
          <pub-id pub-id-type="doi">10.18653/v1/2021.findings-emnlp.59</pub-id>
        </nlm-citation>
      </ref>
      <ref id="ref27">
        <label>27</label>
        <nlm-citation citation-type="journal">
          <person-group person-group-type="author">
            <name name-style="western">
              <surname>Voorhees</surname>
              <given-names>EM</given-names>
            </name>
            <name name-style="western">
              <surname>Tice</surname>
              <given-names>DM</given-names>
            </name>
          </person-group>
          <article-title>The TREC question answering track</article-title>
          <source>Nat Lang Eng</source>
          <year>2002</year>
          <volume>7</volume>
          <issue>4</issue>
          <fpage>361</fpage>
          <lpage>378</lpage>
          <pub-id pub-id-type="doi">10.1017/s1351324901002789</pub-id>
        </nlm-citation>
      </ref>
      <ref id="ref28">
        <label>28</label>
        <nlm-citation citation-type="confproc">
          <person-group person-group-type="author">
            <name name-style="western">
              <surname>Devlin</surname>
              <given-names>J</given-names>
            </name>
            <name name-style="western">
              <surname>Chang</surname>
              <given-names>MW</given-names>
            </name>
            <name name-style="western">
              <surname>Lee</surname>
              <given-names>K</given-names>
            </name>
            <name name-style="western">
              <surname>Toutanova</surname>
              <given-names>K</given-names>
            </name>
          </person-group>
          <article-title>BERT: pre-training of deep bidirectional transformers for language understanding</article-title>
          <year>2018</year>
          <conf-name>Proceedings of NAACL-HLT 2019</conf-name>
          <conf-date>2019 June 2-7</conf-date>
          <conf-loc>Minneapolis, Minnesota</conf-loc>
          <fpage>4171</fpage>
          <lpage>4186</lpage>
          <comment>
            <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://doi.org/10.48550/arXiv.1810.04805"/>
          </comment>
        </nlm-citation>
      </ref>
      <ref id="ref29">
        <label>29</label>
        <nlm-citation citation-type="journal">
          <person-group person-group-type="author">
            <name name-style="western">
              <surname>Jin</surname>
              <given-names>Q</given-names>
            </name>
            <name name-style="western">
              <surname>Kim</surname>
              <given-names>W</given-names>
            </name>
            <name name-style="western">
              <surname>Chen</surname>
              <given-names>Q</given-names>
            </name>
            <name name-style="western">
              <surname>Comeau</surname>
              <given-names>DC</given-names>
            </name>
            <name name-style="western">
              <surname>Yeganova</surname>
              <given-names>L</given-names>
            </name>
            <name name-style="western">
              <surname>Wilbur</surname>
              <given-names>WJ</given-names>
            </name>
            <name name-style="western">
              <surname>Lu</surname>
              <given-names>Z</given-names>
            </name>
          </person-group>
          <article-title>MedCPT: contrastive pre-trained transformers with large-scale PubMed search logs for zero-shot biomedical information retrieval</article-title>
          <source>Bioinformatics</source>
          <year>2023</year>
          <volume>39</volume>
          <issue>11</issue>
          <fpage>btad651</fpage>
          <comment>
            <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://europepmc.org/abstract/MED/37930897"/>
          </comment>
          <pub-id pub-id-type="doi">10.1093/bioinformatics/btad651</pub-id>
          <pub-id pub-id-type="medline">37930897</pub-id>
          <pub-id pub-id-type="pii">7335842</pub-id>
          <pub-id pub-id-type="pmcid">PMC10627406</pub-id>
        </nlm-citation>
      </ref>
      <ref id="ref30">
        <label>30</label>
        <nlm-citation citation-type="journal">
          <person-group person-group-type="author">
            <name name-style="western">
              <surname>Wang</surname>
              <given-names>L</given-names>
            </name>
            <name name-style="western">
              <surname>Yang</surname>
              <given-names>N</given-names>
            </name>
            <name name-style="western">
              <surname>Huang</surname>
              <given-names>X</given-names>
            </name>
            <name name-style="western">
              <surname>Jiao</surname>
              <given-names>B</given-names>
            </name>
            <name name-style="western">
              <surname>Yang</surname>
              <given-names>L</given-names>
            </name>
            <name name-style="western">
              <surname>Jiang</surname>
              <given-names>D</given-names>
            </name>
          </person-group>
          <article-title>Text embeddings by weakly-supervised contrastive pre-training</article-title>
          <source>arXiv</source>
          <comment>Preprint posted online on December 7, 2022</comment>
          <comment>
            <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://doi.org/10.48550/arXiv.2212.03533"/>
          </comment>
        </nlm-citation>
      </ref>
      <ref id="ref31">
        <label>31</label>
        <nlm-citation citation-type="journal">
          <person-group person-group-type="author">
            <name name-style="western">
              <surname>Abdin</surname>
              <given-names>M</given-names>
            </name>
            <name name-style="western">
              <surname>Jacobs</surname>
              <given-names>SA</given-names>
            </name>
            <name name-style="western">
              <surname>Awan</surname>
              <given-names>AA</given-names>
            </name>
            <name name-style="western">
              <surname>Bach</surname>
              <given-names>N</given-names>
            </name>
          </person-group>
          <article-title>Phi-3 technical report: a highly capable language model locally on your phone</article-title>
          <source>arXiv</source>
          <comment>Preprint posted online on April 22, 2024</comment>
          <comment>
            <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://arxiv.org/abs/2404.14219"/>
          </comment>
        </nlm-citation>
      </ref>
      <ref id="ref32">
        <label>32</label>
        <nlm-citation citation-type="confproc">
          <person-group person-group-type="author">
            <name name-style="western">
              <surname>Labrak</surname>
              <given-names>Y</given-names>
            </name>
            <name name-style="western">
              <surname>Bazoge</surname>
              <given-names>A</given-names>
            </name>
            <name name-style="western">
              <surname>Morin</surname>
              <given-names>E</given-names>
            </name>
            <name name-style="western">
              <surname>Gourraud</surname>
              <given-names>P</given-names>
            </name>
            <name name-style="western">
              <surname>Rouvier</surname>
              <given-names>M</given-names>
            </name>
            <name name-style="western">
              <surname>Dufor</surname>
              <given-names>R</given-names>
            </name>
          </person-group>
          <article-title>BioMistral: a collection of open-source pretrained large language models for medical domains</article-title>
          <year>2024</year>
          <conf-name>Findings of the Association for Computational Linguistics: ACL 2024</conf-name>
          <conf-date>August 11-16, 2024</conf-date>
          <conf-loc>Bangkok, Thailand</conf-loc>
          <fpage>5848</fpage>
          <lpage>5864</lpage>
          <pub-id pub-id-type="doi">10.18653/v1/2024.findings-acl.348</pub-id>
        </nlm-citation>
      </ref>
      <ref id="ref33">
        <label>33</label>
        <nlm-citation citation-type="web">
          <source>MTSamples</source>
          <access-date>2026-07-22</access-date>
          <comment>
            <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://www.mtsamples.com/">https://www.mtsamples.com/</ext-link>
          </comment>
        </nlm-citation>
      </ref>
      <ref id="ref34">
        <label>34</label>
        <nlm-citation citation-type="journal">
          <person-group person-group-type="author">
            <name name-style="western">
              <surname>Zhao</surname>
              <given-names>Z</given-names>
            </name>
            <name name-style="western">
              <surname>Jin</surname>
              <given-names>Q</given-names>
            </name>
            <name name-style="western">
              <surname>Chen</surname>
              <given-names>F</given-names>
            </name>
            <name name-style="western">
              <surname>Peng</surname>
              <given-names>T</given-names>
            </name>
            <name name-style="western">
              <surname>Yu</surname>
              <given-names>S</given-names>
            </name>
          </person-group>
          <article-title>A large-scale dataset of patient summaries for retrieval-based clinical decision support systems</article-title>
          <source>Sci Data</source>
          <year>2023</year>
          <month>12</month>
          <day>18</day>
          <volume>10</volume>
          <issue>1</issue>
          <fpage>909</fpage>
          <pub-id pub-id-type="doi">10.1038/s41597-023-02814-8</pub-id>
          <pub-id pub-id-type="medline">38110415</pub-id>
          <pub-id pub-id-type="pii">10.1038/s41597-023-02814-8</pub-id>
          <pub-id pub-id-type="pmcid">PMC10728216</pub-id>
        </nlm-citation>
      </ref>
      <ref id="ref35">
        <label>35</label>
        <nlm-citation citation-type="journal">
          <person-group person-group-type="author">
            <name name-style="western">
              <surname>Khodadad</surname>
              <given-names>M</given-names>
            </name>
            <name name-style="western">
              <surname>Kasmaee</surname>
              <given-names>AS</given-names>
            </name>
            <name name-style="western">
              <surname>Astaraki</surname>
              <given-names>M</given-names>
            </name>
            <name name-style="western">
              <surname>Mahyar</surname>
              <given-names>H</given-names>
            </name>
          </person-group>
          <article-title>Towards domain specification of embedding models in medicine</article-title>
          <source>arXiv</source>
          <comment>Preprint posted online on July 25, 2025</comment>
          <pub-id pub-id-type="doi">10.48550/arXiv.2507.19407</pub-id>
        </nlm-citation>
      </ref>
      <ref id="ref36">
        <label>36</label>
        <nlm-citation citation-type="journal">
          <person-group person-group-type="author">
            <name name-style="western">
              <surname>Nakagawa</surname>
              <given-names>S</given-names>
            </name>
            <name name-style="western">
              <surname>Schielzeth</surname>
              <given-names>H</given-names>
            </name>
          </person-group>
          <article-title>A general and simple method for obtaining R² from generalized linear mixed-effects models</article-title>
          <source>Methods Ecol Evol</source>
          <year>2012</year>
          <month>12</month>
          <day>03</day>
          <volume>4</volume>
          <issue>2</issue>
          <fpage>133</fpage>
          <lpage>142</lpage>
          <pub-id pub-id-type="doi">10.1111/j.2041-210x.2012.00261.x</pub-id>
        </nlm-citation>
      </ref>
      <ref id="ref37">
        <label>37</label>
        <nlm-citation citation-type="web">
          <person-group person-group-type="author">
            <name name-style="western">
              <surname>Mikkelsen</surname>
              <given-names>Y</given-names>
            </name>
          </person-group>
          <article-title>Clinical embedding layer analysis (revision repository)</article-title>
          <source>Zenodo</source>
          <year>2026</year>
          <access-date>2026-07-14</access-date>
          <comment>
            <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://doi.org/10.5281/zenodo.21358857">https://doi.org/10.5281/zenodo.21358857</ext-link>
          </comment>
        </nlm-citation>
      </ref>
      <ref id="ref38">
        <label>38</label>
        <nlm-citation citation-type="web">
          <person-group person-group-type="author">
            <name name-style="western">
              <surname>Mikkelsen</surname>
              <given-names>Y</given-names>
            </name>
          </person-group>
          <article-title>Clinical RAG retrieval benchmark (companion study repository)</article-title>
          <source>Zenodo</source>
          <year>2026</year>
          <access-date>2026-05-30</access-date>
          <comment>
            <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://doi.org/10.5281/zenodo.19482585">https://doi.org/10.5281/zenodo.19482585</ext-link>
          </comment>
        </nlm-citation>
      </ref>
      <ref id="ref39">
        <label>39</label>
        <nlm-citation citation-type="web">
          <person-group person-group-type="author">
            <name name-style="western">
              <surname>Mikkelsen</surname>
              <given-names>Y</given-names>
            </name>
          </person-group>
          <article-title>Clinical-embedding-fix: post-hoc corrections for degenerate clinical text embeddings in RAG pipelines</article-title>
          <source>Zenodo</source>
          <year>2026</year>
          <access-date>2026-05-30</access-date>
          <comment>
            <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://doi.org/10.5281/zenodo.20412005">https://doi.org/10.5281/zenodo.20412005</ext-link>
          </comment>
        </nlm-citation>
      </ref>
    </ref-list>
  </back>
</article>
