Internal Medicine: Distilled

By Omar Nabil Metwally, MD
Internal Medicine Physician
First published 30 August 2026 · Revised and expanded with references


One of the most important classes that I attended as a Clinical Informatics Fellow at UCSF was a course on clinical research. This reflection surprises me in hindsight because I long associated “Clinical Informatics”, which was then a new medical specialty (I was in UCSF’s second class and their third Clinical Informatics Fellow), with the intersection of technology and patient care — a niche for technologically-affine physicians capable of transforming healthcare, if only (so the reasoning) they connected with the right decision makers, got on the right Teams calls, and joined the most elite mailing lists. I was not expecting to sit down almost 10 years later to write a blog post about what I then considered the driest of subjects: biostatistics.

In retrospect, I had a lot to learn, and I could have learned those valuable lessons with a more open and receptive mindset. I consider Professor Pletcher’s course significant in two regards. First, it gave me an opportunity to think deeply about the tests that doctors (especially non-surgeon physicians) routinely order on behalf of their patients and are expected to interpret and act upon correctly. Second, it planted the seed for the realization, many years after my time at Parnassus, that the specialty of Internal Medicine boils down to one concept: a priori likelihood.

There are two types of doctors that take care of people: physicians and surgeons. Physicians — doctors like internists and pediatricians — do a bulk of their work in the cognitive realm, which often takes the form of making a short-list of diagnoses and honing in possibilities (in medical jargon, this is called making a differential diagnosis). This is in contrast to surgeons, whose work tends to be more operative in a hands-on, physical sense than a physician. A physician who does not understand the tests they are ordering, the diagnostic consequences of their choice of test, and how to interpret the results, is like a surgeon without the physical skills necessary to perform a particular operation. I’m deliberately over-simplifying to make the point that despite the fact that test ordering and interpretation has traditionally been a core competency of physicians in general, I encounter an alarming number of physicians, new and experienced alike, who lack the skills to understand what they’re ordering (or not ordering) and how to correctly interpret positive and negative results. Equally alarming, and I suspect related, is organizational policies that narrowly prescribe diagnostic protocols in order to move a patient from point A to point B on the healthcare labyrinth — in the process, stripping physicians of what little autonomy they still have to exercise their clinical judgement.

The objective of this writing is to capture that lesson for any doctor who will ever order (or decide not to order) a test and be expected to interpret the result — if they are lucky to even have the choice.

The order that sounds correct

A middle-aged man presents with months of progressive hip and buttock pain. The pain is worst in the first hours after waking and after prolonged sitting. It wakes him at night. He has a years-long history of intermittent low back pain, a personal history of eosinophilic esophagitis, and a family history of autoimmune disease. Plain films show mild bilateral hip osteoarthritis — premature for his age. He asks for a rheumatology referral.

A reasonable-sounding plan: order an ANA and rheumatoid factor first. If negative, hold the referral. If positive, refer.

That plan has the shape of good medicine — screen, then escalate. It is also, for this presentation, close to backwards. Two errors are stacked inside it, and the second one is the more instructive.

Error one: the tests cannot see the disease

The clinical question here is axial spondyloarthritis. AxSpA is a seronegative spondyloarthropathy — and “seronegative” is not decoration. RF is expected to be negative. ANA has no association with the disease and appears nowhere in the ASAS classification criteria [1]. No result either test returns can move the probability of axSpA in either direction.

So the proposed gate was built from tests that cannot detect the condition being gated. A negative result would have been recorded as reassurance while carrying zero information about the actual question.

Error two: predictive value belongs to the patient, not the test

Here is the concept the wards tend to erode: sensitivity and specificity are properties of the assay. Positive and negative predictive value are properties of the encounter. The same ANA, on the same analyzer, means something entirely different depending on who is in the chair.

Do the arithmetic once and it inoculates you.

Worked example: the ANA in this patient

ANA is positive in roughly 20–30% of healthy adults at a 1:40 titer, and 10–15% at 1:80 [2]. Our patient has monoarticular hip pain with a mechanical mechanism, an axial pain pattern, and no features of connective tissue disease. Be generous and set the pretest probability of an ANA-associated disease at 0.5% — 1 in 200.

Per 1,000 similar patients Test positive Test negative
Disease present (5) 5 0
Disease absent (995) ~120 ~875
Total ~125 ~875
Assuming ~95% sensitivity and ~88% specificity at a 1:80 cutoff [3].

PPV = 5 / 125 = 4%. Twenty-four out of every twenty-five positives are false. A positive ANA in this patient is not a signal — it is noise wearing the costume of data.

And the costume is expensive. A positive ANA in the chart is never inert: it cascades to an ENA panel, a dsDNA, a “possible early connective tissue disease” note, a referral placed for the wrong reason, a patient reading about lupus at 2 a.m., and a label copied forward through every future chart review. The harm of a low-value test is rarely the test. It is the machinery the result switches on.

Worked example: the test that actually earns its place

Now run HLA-B27 through the same patient. Sensitivity is roughly 85% for axSpA — best established for radiographic disease, somewhat lower for non-radiographic axSpA and in non-European ancestries — with specificity around 90%, capped by a background allele frequency of roughly 6–8% in US populations of European descent (and higher still in Scandinavia) [4]. That yields LR+ ≈ 9 and LR− ≈ 0.15.

Given the inflammatory pain pattern, morning predominance, hip involvement, male sex, and family history, a pretest probability of 35% is defensible (pretest odds ≈ 0.54).

  • Positive: 0.54 × 9 = posterior odds 4.9 → ~83% post-test probability. Decision-changing.
  • Negative: 0.54 × 0.15 = posterior odds 0.08 → ~7% post-test probability. Also decision-changing, in the opposite direction.

That is the signature of a test worth ordering: both results move you substantially, from an intermediate starting point. Note the dependency, though — run that same B27 on unselected back pain with a 2% pretest probability and the PPV falls under 20%. The likelihood ratio is a property of the assay; the usefulness is a property of your clinical reasoning.

And the tests that are neither: CRP and ESR

Worth naming because their prominence in “inflammatory workups” oversells them. CRP is elevated in only 40–50% of axSpA — a normal value carries a likelihood ratio near 1 and excludes essentially nothing. Their value here is prognostic (elevated CRP is one of the strongest predictors of radiographic progression, alongside baseline syndesmophytes and smoking [5]), as a baseline for treatment monitoring and ASDAS scoring, and for flagging infection. Diagnostically, they are close to inert. Order them knowing which job they are doing.

A worked example everyone lived through: the COVID home test

During the pandemic the federal government mailed free rapid antigen tests to any household that asked — well over a billion tests moved through federal programs before distribution wound down and was finally suspended in March 2025 [6]. Today the same tests sit on pharmacy shelves at $10–20 a box, and most people with a cough don’t bother. That shift is the thesis of this essay playing out at national scale — though the arithmetic contains a surprise worth sitting with.

At the Omicron peak, roughly half of symptomatic adults with acute respiratory symptoms actually had COVID: in CDC’s outpatient surveillance network, 56% of symptomatic adults tested positive during Omicron predominance [7]. Rapid antigen tests in symptomatic people carry a pooled sensitivity of about 73% and a specificity of about 99.6% [8]. Per 1,000 symptomatic people at 50% prevalence: 365 true positives against 2 false ones — a positive predictive value near 99.5%. The test was superb, and it was superb because of who was taking it.

Today, the same symptoms in the same person carry perhaps a 10% probability of COVID in a typical week — the rest is rhinovirus, RSV, influenza, and everything else — and closer to 3% in a deep trough between waves. Same assay, same operating characteristics. At 10% prevalence: 73 true positives against about 4 false ones — PPV around 95%. At 3%: PPV around 85%.

Here is where the intuition usually goes wrong, and it is the most instructive moment in this essay. The PPV fell — but nowhere near as far as it did for the ANA in our patient, whose PPV was 4% at low pretest probability. Why the difference? Specificity. The antigen test generates false positives at a rate of about 0.4%; the ANA at a 1:80 cutoff generates them at about 12%. When prevalence drops, a test with a fraction-of-a-percent false-positive rate degrades gracefully, and a test with a double-digit false-positive rate falls off a cliff. Pretest probability sets your ceiling; specificity determines how fast you fall from it. A positive home COVID test today is still probably real. A positive ANA in a patient without connective-tissue-disease features almost never is.

One more property worth internalizing: the sensitivity of a strategy is not the sensitivity of a test. A single antigen test on the first day of symptoms detects only about 60% of infections, but two tests taken 48 hours apart detect over 93% [9] — which is exactly why the FDA recommends repeat testing after a negative result [10]. Operating characteristics belong not just to the assay but to how you deploy it.

So why did universal home testing wind down? Partly for structural reasons that have nothing to do with epidemiology: the public health emergency ended in May 2023 and took the insurer free-test mandate with it, federal distribution stopped, and the tests now cost real money. But the clinical logic tracks the third question of this essay: will the result change what I do next? Under CDC’s unified respiratory-virus guidance, isolation is now symptom-based — stay home until improving and fever-free for 24 hours — regardless of which virus you have [11]. For a healthy, low-risk adult with mild symptoms, a positive test changes little about their own care.

And notice that the answer flips right back for the people in whom question three is still answered yes: anyone eligible for antiviral therapy, where a positive result starts a five-day treatment clock [12]; anyone about to sit down with an immunocompromised relative; healthcare workers and congregate settings, which the relaxed guidance explicitly does not cover. The newer over-the-counter combination flu/COVID tests sharpen the point further — in a high-risk patient, distinguishing influenza from COVID selects between two different antivirals. Same test, same prevalence, opposite recommendation — because the question being asked is different.

We ran a natural experiment on 300 million people, and the lesson was the one Professor Pletcher was teaching in a classroom on Parnassus: the test never changed. Who we gave it to, and what we planned to do with the answer, changed everything.

The three questions that replace the panel

The fix isn’t memorizing which panel goes with which complaint — panels are how we got here. It’s a habit, asked in order, before anything is ordered:

  1. What disease am I testing for? Named and specific, out loud. Not “inflammatory something.” If you can’t name the disease, you can’t choose the test — tests are only interpretable against a named hypothesis.
  2. What is my pretest probability — even crudely? “Under 5%, coin flip, or over 50%” is enough resolution to change behavior. The discipline is committing to a number before the result exists. A probability estimated afterward isn’t an estimate; it’s a rationalization.
  3. Will either result change what I do next? If positive changes nothing and negative changes nothing, the test’s only outputs are cost, noise, and cascade risk.

Run the case through them. Disease: axSpA. Pretest probability: intermediate. Will ANA/RF change the next step? No — the MRI and the referral were indicated regardless of serology. Question three deletes the ANA in about four seconds. It also reveals something sharper: the problem wasn’t only which tests formed the gate. There should have been no gate.

Why good physicians make this error

This is not stupidity. It is a systems failure with a cognitive assist, and you should recognize the ingredients, because they will be your working conditions:

  • Order sets are frozen decisions. Someone decided once that ANA + RF = “rheum labs,” and the EHR has been re-making that decision on autopilot ever since. Every checkbox panel is a colleague from the past overriding your present reasoning — sometimes correctly, and you won’t know which times unless you look.
  • Covering is medicine at low resolution. An unfamiliar chart, an intermediary, a message queue. Pattern-matching to “joint pain → rheum panel” is what cognition produces under those constraints. The antidote isn’t heroic effort; it’s cheap heuristics that survive low resolution — which is what the three questions are.
  • “Baseline” is a thought-terminating word. Nobody argues with a baseline. But a baseline is just a test whose pretest probability nobody bothered to state.

And the sequel matters. When this reasoning was raised — politely, in writing — the physician reviewed the chart and ordered the appropriate labs within two hours. That responsiveness is the trait to emulate. The initial error was ordinary; the rapid update was excellent medicine. If you take a villain from this story, you’ve misread it. The villain is the reflex, and every one of us has it.

The distillation

A test result is not information about the patient. It is information about the patient conditional on why you ordered it.

The test doesn’t know who you ordered it on. You do. The entire interpretive weight of every result you will ever receive rests on the probability you assigned — explicitly, or by default — before you clicked the order.

Assign it explicitly. Every time. Ten seconds, and it is the highest-yield ten seconds in diagnostic medicine.


References

  1. Rudwaleit M, van der Heijde D, Landewé R, et al. The development of Assessment of SpondyloArthritis international Society classification criteria for axial spondyloarthritis (part II): validation and final selection. Ann Rheum Dis. 2009;68(6):777–783.
  2. Tan EM, Feltkamp TE, Smolen JS, et al. Range of antinuclear antibodies in “healthy” individuals. Arthritis Rheum. 1997;40(9):1601–1611.
  3. Leuchten N, Hoyer A, Brinks R, et al. Performance of antinuclear antibodies for classifying systemic lupus erythematosus: a systematic literature review and meta-regression of diagnostic data. Arthritis Care Res (Hoboken). 2018;70(3):428–438.
  4. Reveille JD, Hirsch R, Dillon CF, Carroll MD, Weisman MH. The prevalence of HLA-B27 in the United States: data from the US National Health and Nutrition Examination Survey, 2009. Arthritis Rheum. 2012;64(5):1407–1411.
  5. Poddubnyy D, Rudwaleit M, Haibel H, et al. Rates and predictors of radiographic sacroiliitis progression over 2 years in patients with axial spondyloarthritis. Ann Rheum Dis. 2011;70(8):1369–1374.
  6. US Department of Health and Human Services. Federal at-home COVID-19 test distribution program (COVIDTests.gov), January 2022; distribution suspended March 2025.
  7. Kim SS, Chung JR, Talbot HK, et al. Effectiveness of two and three mRNA COVID-19 vaccine doses against Omicron- and Delta-related outpatient illness among adults, October 2021–February 2022. Influenza Other Respir Viruses. 2022;16(6):975–985.
  8. Dinnes J, Sharma P, Berhane S, et al. Rapid, point-of-care antigen tests for diagnosis of SARS-CoV-2 infection. Cochrane Database Syst Rev. 2025;CD013705.
  9. Soni A, Herbert C, Lin H, et al. Performance of rapid antigen tests to detect symptomatic and asymptomatic SARS-CoV-2 infection: a prospective cohort study. Ann Intern Med. 2023;176(7):975–982.
  10. US Food and Drug Administration. At-home COVID-19 antigen tests: take steps to reduce your risk of false negative results. FDA Safety Communication, August 2022.
  11. Centers for Disease Control and Prevention. Respiratory virus guidance. March 2024.
  12. US Food and Drug Administration. Nirmatrelvir-ritonavir (Paxlovid) prescribing information: treatment of mild-to-moderate COVID-19 in patients at high risk for progression, initiated within 5 days of symptom onset.

Clinical details are composite and deidentified. Test characteristics are approximate, vary by assay, cutoff, population, and timing, and are cited to their principal sources above — the arithmetic is illustrative, and the habit is the point.