Draft · not listed anywhere
Going from codes to concepts: building an ontology in healthcare
Tarush Aggarwal · August 2026 · 11 min
- Sector
- Healthcare, patient intake
- Region
- New York, 1,000+ clinics
- Company
- Intake software for US practices
- Before
- Four codes for one disease
- After
- 105,000 conditions, one concept
- Unlocked
- One answer per question
TL;DR
Ask a healthcare company how many of its patients have diabetes and you will not get one answer. You will get one per system, because the same disease arrives wearing a different code from every EMR and all of those codes are correct.
A semantic layer cannot close that gap. It names columns. It has no way to know that one code is a kind of another. In healthcare that hierarchy is the answer, which makes an ontology the layer everything else has to sit on.
Most data work ends at a semantic layer, and for most companies that is the right place to stop. A semantic layer is an agreement about names. This column is revenue. That table is a customer. Here is how you join them. It holds up because your company invented the things being named. You decided what a customer is, so you write it down once and move on.
Healthcare never got that. The things being named were defined outside the company, by standards bodies, over about sixty years, and almost nothing was ever retired.
Four codes, one disease
A practice on Epic sends E11.9. A legacy system sends 250.00. An athenahealth practice sends 44054006. A patient filling in a form types "diabetes, type 2". Four encodings of one disease, and every one of them is a correct clinical record.
Every healthcare data team papers over this the same way: a list of codes that count as diabetes, kept in a spreadsheet, owned by whoever built the first report. That list is wrong the day a new code ships, and nobody finds out, because the number it produces still looks like a number.
A list cannot hold this because the codes are related to each other, and those relationships are what the question is actually asking about. Type 2 diabetes is a kind of diabetes. Stage 3 chronic kidney disease is a kind of chronic kidney disease. Allergic asthma is a kind of asthma. "How many patients have any diabetes" is a question about a hierarchy, and a flat list has no hierarchy in it.
What an ontology is
Two things: a set of concepts, and the is-a relationships between them. Diabetes mellitus is a concept. Type 2 diabetes mellitus is a concept and a child of the first. Every code any system sends resolves to one concept, and every question walks the relationships instead of matching raw codes.
That buys three things a semantic layer cannot give you:
- One answer per question, whichever system recorded it.
- A hierarchy you can walk. Ask about diabetes, get every subtype under it. Ask about type 2, get that branch only.
- Somewhere to put uncertainty. A code that cannot be resolved lands in a queue instead of being guessed at or dropped.
The part worth taking away
Conversational AI over a warehouse works well in most industries because the semantic layer underneath carries enough meaning to answer with. Healthcare is where that runs out. Point an agent at the warehouse and ask how many patients are diabetic, and it will write a query against whatever codes it can see and hand back a number that is confidently short. Nobody catches it. The number looks fine.
So the ontology is the foundation, and it goes in first. Build it and the reporting layer, the product features and the conversational layer are all correct for the same reason. Skip it and you have shipped a very fluent interface to the wrong answer.
What we built
A New York healthcare company runs the patient intake layer for medical practices across the US: forms, check-in, insurance verification, scheduling. Their software sits in more than a thousand clinics, which means the same handful of diseases reaches them in the dialect of every EMR their customers run, plus whatever patients type into a form themselves.
We do not name clients in this newsletter, so the screenshots below have their branding taken out. Everything else on them is the running service.
We built a proof of concept in a week. It is a standalone service rather than a report or a warehouse model, and it does one job: any code, clinical term or patient answer in, one standard concept out.
The vocabulary layer is OMOP's open clinical vocabularies, SNOMED CT, ICD-10-CM and ICD-9-CM, which carry no license fee. Commercial terminology vendors sell this as a subscription running into six figures a year. The curated content they maintain on top is real work and worth paying for. The base mapping underneath it is not the part you need to buy.
105,000 conditions already resolve. Diabetes is only what the demo walks through.
Everything below is the running service. The patients are synthetic.
1. The form the patient sees
The front door has no codes anywhere near it. Plain questions, plain options, in the words a person would use about their own body.

Four things happen underneath that are worth pulling out.
Comorbidity is the normal case. Diabetes and kidney disease together is what real patients look like, so each condition selected opens its own follow-up and carries its own onset year. People are rarely diagnosed with two things in the same year, and a shared date would be a tidy fiction.
One submit, one record per condition. The whole form posts as a unit and writes a separate clinical record for each condition, which is how an EMR stores them. It has to see the whole answer at once, because "select all that apply" is also a statement about everything the patient did not select.
"Not sure" is a real answer. A patient who knows they have diabetes but not which type gets stored against the parent concept and flagged, rather than having a subtype invented for them. They still count correctly in every rollup that asks about diabetes. This is the thing flat code lists cannot do at all: there is no code for "diabetes, type unknown" that also behaves like diabetes in a report.
A patient's answer never overwrites a clinician's. Re-submitting updates that patient's own reported answer in place. Records that arrived from an EMR are left alone. If Epic recorded diabetes, she stays in that cohort on the strength of Epic's record, and reconciling the two observers is a review problem rather than something the write path should decide on its own.
2. Every dialect, one concept
The translator is the engine the rest of it calls. Give it a code and a vocabulary and it resolves through OMOP's mappings.

E11.9, 250.00 and 44054006 all land on the same concept. That is the easy half.
The half that matters is what it does when it cannot be sure. Type IDDM, the old abbreviation for insulin-dependent diabetes, and the service refuses to resolve it. The reason is on screen: the highest-ranked hit matches a synonym containing NIDDM, which is the clinical opposite, and substring ranking has no way to know that. So it stores concept_id 0 and queues the term for a person.
Anyone demonstrating 100% automatic free-text mapping is demonstrating silent errors. You will not see them, which is the problem with them.
The same write path runs the nightly feed. Sixteen records from three EMRs in their native dialects, fourteen mapped automatically, two local codes that exist only inside the sending system dropped into the review queue. Run the file again and nothing duplicates.
That last bit is not a detail. Feeds get retried and files get replayed, and a patient acquiring a second copy of their diagnoses every time is how a clean system quietly stops being one.
The queue is where the honest part lives. Each item carries the raw code, the system it came from, the patient it was going to be attached to, and suggested concepts where the code has a name to search on. A local code with no name in any vocabulary gets no suggestion at all, and says so. When a reviewer approves a mapping it is written to disk, consulted by every later lookup, and backfilled onto the patients who arrived while the code was still unknown. Nobody has to remember to re-run anything.
3. Reporting that survives the encodings
This is the question the whole thing exists to answer, and now it answers once.

43 patients are recorded with some form of diabetes, across four source systems and 22 distinct encodings. A report filtering on E11 would have missed most of them, and would have looked perfectly healthy while doing it.
Underneath, the rollup is one indexed query against a precomputed ancestor table, so it behaves the same at four concepts or ten million. Change the filter to type 2 only and the count follows the tree. Change it to hypertension and the same engine answers on a different branch, because nothing in it is diabetes-specific.
Two things on that screen are there because leaving them off would be dishonest.
The first is the phrasing. It says patients recorded with a diabetes diagnosis, never patients who have diabetes. Those are different claims and only one of them is supported by the data.
The second is the caveat in coral. Three patients arrived with codes that map to a diabetic complication, and SNOMED files a complication under the body system it affects rather than the disease that caused it. So those three are not descendants of diabetes mellitus, and this rollup does not find them. They are excluded, not lost, and switching the filter to any recorded condition shows them. That is the vocabulary's modelling rather than a defect in the query, and closing it properly means curated condition sets, which is exactly the content work a terminology subscription's fee actually buys.
A system that surfaces its own edges is one you can plan around. A system that quietly returns 43 and says nothing is one you will trust for a year and then stop trusting entirely.
4. The graph, up and down
Every rollup above is walking a structure. This draws it.

Conditions are nodes, is-a relationships are edges, and every node carries a live patient count. You can start at all four condition families and go down, or start at diabetes mellitus and walk its subtypes. Click any node and you land on that cohort in the report.
Two properties make it trustworthy rather than decorative.
Every node's number is the same query the reporting screen runs, and a test walks the whole graph asserting the two agree. A diagram that disagreed with the report it is supposed to explain would be worse than having no diagram.
And it tells you what it is not showing. Hypertensive disorder has 33 children in SNOMED and the screen draws 8 of them, ranked by patient count, with a dashed stub saying how many were left out. A picture that quietly showed 8 of 33 would be the same lie as an API that did.
This is the view people mean when they say Palantir-style ontology. The structure was always in the vocabulary. Drawing it is what makes it usable by a human, and having it queryable is what makes it usable by an agent.
What it cost to find out
A week, one engineer, and no license fees. That is the point of doing it this way. The expensive version of this decision is a nine month build against a vendor contract you signed before you knew whether the shape was right.
The proof of concept answers the questions that actually change the design. Should this be the system of record for conditions, or a translation service other systems call? Does it get embedded in the product, or exposed as an API? Which standards do you actually need? How many years of existing records get normalized at go-live? Those are much easier to answer with something running in front of you.
FAQ
Isn't this what a terminology vendor already sells?
Partly, and the part they sell that is genuinely hard is the curated content: maintained lists of every code that should count toward a cohort, complications included. The base mapping between standards is published free in OMOP. Buy the content if you need it. You do not need to rent the plumbing, and renting it means your patient conditions live inside someone else's API on a renewal cycle.
Why not just let a model do the mapping?
Because a language model will always give you an answer, and on a diagnosis a confident near miss is worse than a blank. The IDDM case is the whole argument in one lookup: the most plausible-looking match resolves to the clinical opposite of the right one. Models are useful here for ranking suggestions in the review queue, where a person makes the call.
Does this replace the EMR?
No. It sits beside the systems that already hold the records and gives them one shared vocabulary. Whether it becomes the system of record for conditions or stays a translation service is a real decision with real consequences, and it is worth making deliberately rather than by accident.
Does this only work for healthcare?
The pattern needs an industry where the vocabulary was defined outside your company, is hierarchical, and changes without asking you. Healthcare is the clearest case. Insurance, logistics, manufacturing parts catalogs and financial instrument taxonomies all have a version of it.
Tarush
Keep reading
How hedge funds are building a context layer around Excel
Why Excel, documents and email resist AI, how a hedge fund kept every spreadsheet and got a data platform anyway, and what becomes possible once the data sits underneath.
Build vs buy: How AI is changing how we think about managing risk
Why buying off the shelf was always a risk decision, what changes when building gets cheap, and the three ways of implementing AI in your organisation.
Going from trust to proof: How AI is changing operations
Why operations scale by hiring people to check people, how a photo now proves a class happened, and what automated itself once the records were clean.
