02-1. Essay

Okay, Then How Does AI Learn an Industry?

Healthcare, legal, telecom. Very different industries, yet the same three questions keep showing up when we try to build domain-specific AI.

의료, 법률, 통신. 전혀 다른 산업이지만 산업 특화 AI를 만들기 위해 반복되는 세 가지 질문들.

Look across AI in healthcare, legal, and telecom, and one question starts to emerge: does teaching AI an industry simply mean training it on more industry data?

In this piece, I followed real examples from MedGemma, Abridge, Harvey, CoCounsel, and the Telco LLM work I was directly involved in to understand what domain-specific AI actually needs to learn. Along the way, the same three questions kept resurfacing: understand the current situation, find the knowledge needed to solve it, and follow the way an expert actually works. I wanted to use those three questions to understand how Domain Context gets built.

In the last piece, I looked at how the exams we give AI are changing.

We are moving from testing whether a model can pick the right answer to testing whether it can carry an expert-level project all the way from start to finish.

It feels a little like moving from standardized tests to a hiring assessment, or maybe even a promotion review.

What made Agents’ Last Exam especially interesting was that even SOTA models had a 0% pass rate. Two of the biggest failure modes were Wrong Strategy and Domain Knowledge Gap.

In other words, if AI is going to work like an expert, it has to understand the industry and know what actually matters in the situation in front of it.

Which led me to the question for this piece:

How do we actually give AI domain expertise?

First, I Split the Industries into Two Groups

Before diving into individual industries, I had to decide where to start. I ended up grouping them into two broad categories.

  • Group A. Healthcare, legal, telecom
  • Group B. Manufacturing, semiconductors, defense

The first group is heavily driven by specialized knowledge and data. Much of that information is also sensitive and shaped by regulation, security, and privacy constraints.

The second group sits one step closer to the physical world. Here, AI has to understand the current state of a machine on a factory floor, a semiconductor design, or a sensor environment on a battlefield, then make decisions and act within that physical system. This is exactly where the conversation starts to overlap with one of the hottest areas in AI today: Physical AI.

In 02-1, I am starting with healthcare, legal and telecom. Rather than beginning with a framework and forcing examples into it, I wanted to follow what the major players are actually building, alongside recent research, and see what patterns emerge across industries.

Then in 02-2, I will take the same questions into manufacturing, semiconductors and defense.

Healthcare first.

Healthcare

Knowing medicine and understanding this patient are not the same thing

The most intuitive way to specialize healthcare AI is to teach the model more medicine.

Google’s MedGemma 1.5 is a good example. It specializes the broader Gemma family for medical text and imaging, extending its capabilities into 3D CT and MRI, pathology images and EHR understanding. In other words, one approach is simply to make the model itself a better medical model.

The MedGemma collection, spanning 2D imaging, text, speech and advanced imaging
The range of information medical AI has to read keeps widening. MedGemma 1.5 reaches past text and 2D images into CT, MRI, pathology slides and the changes across a patient’s prior imaging.Source: Google Research, “Next-generation medical image interpretation with MedGemma 1.5” (2026)

So if we make an AI read enough medical textbooks, does it eventually become a doctor?

Looking across the major healthcare AI updates of 2026, the answer seems to be: not quite.

OpenAI launched ChatGPT for Healthcare in January 2026, then expanded Health in ChatGPT in July so users could connect medical records and Apple Health.

Health in ChatGPT screens for connecting medical records and Apple Health
What does an AI that knows your health actually look like? Health in ChatGPT connects medical records and Apple Health so a user’s question can be understood inside the context of their real health information.Source: OpenAI, “Health in ChatGPT” (2026)

Model specialization is clearly part of the story. But at the same time, major players are building something around the model: access to the patient’s actual context.

Different products, different architectures. But when I looked at what they were trying to solve, three questions appeared surprisingly quickly.

① What is happening right now? Patient State

A patient says:

“I’ve been getting really short of breath lately.”

A general-purpose model can list dozens of possible causes. But that is not what a physician actually needs. The real question is why this particular patient is short of breath. You need age, medical history, medications, prior visits, recent lab results, imaging and more.

The words “I’m short of breath” may be identical, but they mean very different things coming from a 20-year-old athlete and a 70-year-old cardiac patient.

Abridge connects the current clinical conversation with ER records, nursing assessments, test results, imaging, and previous visits to build a picture of the patient before the encounter even begins. OpenAI and Anthropic are likewise connecting health records and personal health data so the model can reason within the context of the individual patient.

So in healthcare, answering “What is happening right now?” depends on how well AI can reconstruct the patient’s current state. A lot of the product competition is really about how to turn that patient state into data the model can access at exactly the right moment.

② What knowledge is needed to solve the problem? Clinical Knowledge

Understanding the patient is not enough. The system then has to identify which medical knowledge is actually relevant to the problem in front of it.

Abridge has expanded its evidence sources beyond UpToDate to include NEJM, JAMA, the American Heart Association, and others. In 2026, it has also been moving toward incorporating hospital-specific guidelines and care pathways.

ChatGPT for Healthcare is designed to search peer-reviewed research, public-health guidance and clinical guidelines with citations, while also connecting to an institution’s approved care pathways and operating policies.

The point is not to shovel as much medical knowledge as possible into the model. The harder part is understanding the patient first, then retrieving the evidence that matters for this problem right now.

You do not need to fling the entire medical encyclopedia open. You need to land on the right page for this patient.

③ How does an expert handle it? Care Workflow

The final question is how the work actually gets done. Physicians do not simply read information and generate answers.

Before a visit, they review the patient’s history. During the visit, they interpret symptoms and evidence. Afterward come tests, prescriptions, documentation, follow-up and coordination.

That is why Abridge has expanded its product across the clinical journey, from pre-visit preparation to support during the encounter and post-visit documentation and follow-up. OpenAI is turning recurring clinical tasks such as referrals, prior authorization and patient communication into reusable templates, while Anthropic has introduced healthcare skills that can work across patient records, insurance criteria and clinical guidelines.

So the first pattern I saw in healthcare looked like this:

Patient StateClinical KnowledgeCare Workflow

Understand what is happening to this patient. Retrieve the knowledge needed to solve the problem. Then use it inside the way clinical work actually gets done.

Healthcare seems to struggle most with that third question.

A lot of clinical expertise lives in tacit knowledge, the judgment physicians accumulate through experience but rarely write down in full. Abridge and OpenAI still keep the final judgment with clinicians, and there remains a meaningful gap between AI that can organize information well and AI that can reliably make clinical decisions.

Legal

Then the same structure started showing up again

Next I looked at recent updates from Harvey, Thomson Reuters and LexisNexis. At first, legal felt like a completely different world.

Once I started looking at the products, though, almost the same questions came back.

① What is happening right now? Matter State

If healthcare needs to understand the patient’s current state, legal needs to understand the current state of the matter or transaction.

Hand an M&A lawyer a single contract and ask, “What is risky for our client?” and the document alone is rarely enough. What transaction is this? Who is the client? Which jurisdiction applies? What changed from the previous draft? What positions have already been exchanged during negotiation?

Harvey is moving beyond asking questions of one document at a time. Its agents can connect to multiple sources and a firm’s document management systems, allowing them to work across those materials within a single matter context.

Harvey Agents working through a due-diligence request list across a firm’s documents
Harvey Agents connects document systems such as iManage together with the rest of a matter’s material, then works through what it needs inside that single matter context.Source: Harvey, “Agents”

Healthcare builds the patient state. Legal builds the matter state.

Swap the patient for a legal matter, and apparently every industry is still fighting the same battle: figuring out what is actually going on right now.

② What knowledge is needed to solve the problem? Legal Knowledge and Firm Know-How

The second question looks familiar too. To solve the matter, you first need to know which laws and precedents apply.

Thomson Reuters’ Westlaw is one of the major professional databases lawyers use to research statutes and case law. The next generation of CoCounsel connects AI directly with these verified legal sources. The point is to ground the work in the sources lawyers actually rely on, rather than letting the model hunt around the open internet for something that merely sounds legal. LexisNexis is moving in a similar direction by connecting its own legal databases with AI.

CoCounsel’s verification view, pairing each assertion with the supporting statutory passage
In legal work, what you relied on matters more than how convincing the answer sounds. CoCounsel is strengthening verification so a lawyer can check each assertion back against the legal source it came from.Source: Thomson Reuters, “CoCounsel Legal June 2026 releases”

But legal adds another layer.

Two stacked layers: public legal knowledge underneath, a firm’s institutional knowledge above it
Statutes, regulations and case law form the floor everyone stands on. What separates one firm’s advice from another is the layer above it: past matters, preferred language, internal standards and accumulated judgment.

Two M&A transactions can involve similar law and still produce different advice depending on the client’s risk tolerance, what they refuse to give up in negotiation, or what language the firm has accepted in comparable deals before.

That is why Harvey and DeepJudge introduced the idea of Institutional Intelligence in 2026, connecting a firm’s past work product and accumulated judgment to current AI workflows.

Harvey’s announcement of its Institutional Intelligence partnership with DeepJudge
The statute book is open to everyone. A firm’s experience is not. Harvey and DeepJudge are trying to bring the judgment and know-how accumulated across past matters into current AI work.Source: Harvey, “DeepJudge and Harvey partner to power AI agents with Institutional Intelligence” (2026)

Getting the Westlaw password does not automatically download twenty years of partner experience into your brain.

Legal AI, then, needs to go beyond “knowing a lot of law.” It has to find the law that matters for this matter, and understand how this organization would typically think about it.

③ How does an expert handle it? Legal Workflow

One of the clearest changes in legal AI in 2026 has been the shift from answering questions to performing work.

Harvey introduced Agent Builder so teams and practice areas can turn their own templates and internal processes into reusable agents. An M&A team, for example, may not want to ask from scratch every time: “Please review this contract.” Instead, the workflow might look like this:

Review contractExtract key risk clausesCompare against standard languageDraft revisionsFlag issues requiring human review

That turns the team’s actual way of working into a repeatable process.

LexisNexis has also added pre-built workflows and a Custom Workflow Builder to Protégé so firms can turn their own working methods into repeatable AI processes.

A table of Protégé’s pre-built workflows for small law firms and in-house legal teams
From AI that answers one question to AI that carries out work in order. Protégé is moving a firm’s actual methods and standards into workflows it can run again and again.Source: LexisNexis, “Global launch of Protégé AI workflows for legal professionals” (2026)

Think of it as trying to close the gap between tossing a contract into ChatGPT and actually assigning the work to a junior lawyer.

So the emerging pattern in legal looks like this:

Matter StateLegal + Institutional KnowledgeLegal Workflow

Which looks remarkably similar to healthcare:

Patient StateClinical KnowledgeCare Workflow

See the resemblance?

By this point, the pattern from healthcare was starting to look less like a coincidence.

Telecom

Telecom AI kept landing on the same three questions too

Telecom is a little more personal for me. I spent seven and a half years in the industry.

I was directly involved from the early launch of the Global Telco AI Alliance, or GTAA, through the co-development and go-to-market work for the Telco LLM, and later the formation of Syntelligence AI.

The original GTAA idea was ambitious. SK Telecom in Korea, Deutsche Telekom in Europe, e& in the Middle East, Singtel in Southeast Asia and SoftBank in Japan would jointly develop multilingual Telco LLMs for customer-service use cases.

I still remember the tension in the room when five telecom operators from four continents sat down at the same table for the first time.

Everything was a first, so naturally the mood had all the solemnity of a G20 summit.

Executives from five telecom operators at the Global Telco AI Alliance inaugural summit
MWC Barcelona, 2024. Five operators with different markets and different languages sat down at one table for the first time to build a single Telco LLM.Source: SK Telecom Newsroom (2024)

At the time, I thought that building an LLM that genuinely understood telecom would solve a large part of the problem.

The closer we got to actual customer-service AI, the clearer it became that knowing telecom and doing telecom work were not the same thing.

After enough marathon debates among five operators, almost every discussion seemed to collapse back into the same three questions.

① What is happening right now? Subscriber / Network State

A customer says:

“My Wi-Fi internet isn’t working at home.”

Recognizing that this is a connectivity issue is not particularly hard. The real work starts after that.

The same “my Wi-Fi isn’t working” could point to a device setting, the router, the broadband line entering the home, or the operator’s network itself.

In telecom, understanding this current state means looking beyond subscriber information to the real-time status of the network and devices as well.

That is why, in GTAA, the goal was not to somehow pre-train every real-time state into the model. We needed the AI to pull the latest information from each operator’s systems at the moment a problem occurred and interpret it in context.

The catch was that all five operators had different systems.

So we first defined the core information the AI would need to make a decision in a common way: customer intent, subscribed products, line status, device and equipment information, and current network state. Then we mapped semantically equivalent data from each operator’s subscriber systems and network operations systems into that shared schema.

In other words, we were aligning the meaning of five operators’ data so the AI could read it as one common language.

② What knowledge is needed to solve the problem? Telecom and Operator Knowledge

Once we could represent subscriber and network state in a common format, the next question arrived almost immediately:

“How much can we make common, and where do we have to become operator-specific?”

Concepts such as internet outages, billing disputes or network problems can share a common taxonomy. But the actual product structure, equipment, network environment, policy exceptions and escalation rules vary by operator. So we needed shared Telecom Intelligence, with operator-specific knowledge and policy layered on top.

TeleCom-Bench, released in 2026, reflects a similar distinction. It evaluates not just telecom fundamentals and 3GPP knowledge, but also proprietary product knowledge and real operational workflows.

The TeleCom-Bench framework, split between comprehension and application
TeleCom-Bench splits the evaluation in two: understanding telecom knowledge, and applying that knowledge to real operational work.Source: TeleCom-Bench (2026)

The question is moving from “Does the model know 5G?” to “Can it retrieve and apply the right telecom knowledge to solve the problem happening right now?”

I remember the joy of hearing that an exam would be open-book, followed shortly by the discovery that an open book is surprisingly useless when you have no idea where to look.

③ How does an expert handle it? Operational Workflow

This was the part I personally found hardest.

After five operators had spent more late nights than I care to remember agreeing on a common data schema that could represent subscriber and network state across all five systems, I remember thinking:

Okay… we did it. We’re done. We have actually made it.

Then came the real challenge.

Wait… how exactly are we going to encode the decision logic?

Telecom experts carry around a lot of decision logic that sounds like this:

  • “If the other devices are also down, check the router or the line before blaming the handset.”
  • “If the router looks normal but the line status is abnormal, investigate the network.”
  • “If remote troubleshooting fails, escalate to the next support level.”

TeleCom-Bench makes the gap especially visible. Current models score above 90% on tasks such as intent recognition and key information extraction, but performance drops dramatically when they have to generate an actual solution. Even the best model reached only around 30% on Solution Generation. The researchers call this the Execution Wall.

TeleCom-Bench results table, with solution generation scores far below the rest
Understanding a problem and actually solving it turned out to be different things. On TeleCom-Bench, models cleared 90% on intent recognition but fell to roughly 30% once they had to generate a real solution.Source: TeleCom-Bench (2026)

In retrospect, that is exactly what we were wrestling with too:

Subscriber / Network StateTelecom + Operator KnowledgeOperational Workflow

We were trying to answer the same three questions, one layer at a time.

The GTAA joint venture has since become Syntelligence AI, and that evolution reflects the same shift. Today, Syntelligence focuses on production-ready AI for real telecom operations, including network intelligence, communications trust, and customer experience.

By the Third Time, It Stopped Looking Like a Coincidence

The data is different. The rules of each industry are different. The actual work is completely different. And yet all three industries seem to be answering the same questions.

① What is happening right now?

  • Healthcare: Patient State
  • Legal: Matter State
  • Telecom: Subscriber / Network State

② What knowledge is needed to solve the problem?

  • Healthcare: Clinical Knowledge
  • Legal: Legal + Institutional Knowledge
  • Telecom: Telecom + Operator Knowledge

③ How does an expert handle it?

  • Healthcare: Care Workflow
  • Legal: Legal Workflow
  • Telecom: Operational Workflow

Different industries, but somehow we keep ending up with the same questions.

Can you see the pattern?

Healthcare: Patient StateClinical KnowledgeCare Workflow

Legal: Matter StateLegal + Institutional KnowledgeLegal Workflow

Telecom: Subscriber / Network StateTelecom + Operator KnowledgeOperational Workflow

Seeing the same answers recur across industries made me think of them as the core structure of what I call the Domain Context Layer.

The Domain Context Layer across healthcare, legal and telecom
Different industries, repeating structure. Understand the patient in healthcare, the matter in legal, the subscriber and network in telecom. Then retrieve the knowledge that applies, and finally follow the expert’s workflow.

My Read | Domain-Specific AI May Ultimately Be a Battle Over Who Owns the Context

The simplest version of domain-specific AI goes something like this. Feed it healthcare data, get healthcare AI. Feed it legal data, get legal AI. Feed it telecom data, get telecom AI.

But the products I looked at suggest something more complicated. To build a Domain Context Layer, the system has to answer three questions, and the ingredients used to answer them are completely different across industries.

Healthcare uses patient records, tests and clinical evidence. Legal uses matter data, authoritative legal sources and the firm’s accumulated judgment. Telecom uses subscriber and network state, along with operator-specific products and policies.

Different data. Different systems. Different regulation. Different workflows. And yet the way domain expertise is being assembled looks surprisingly similar.

The important thing may not be stuffing an entire industry into the model. It may be reconstructing, around the model, the context an expert actually uses to make a decision.

Think about how experts work. A physician focuses on the evidence relevant to this patient. A lawyer retrieves the law that matters for this matter, along with the firm’s prior judgment. A telecom expert looks at this subscriber, this network state, and chooses the policy and troubleshooting path that fits.

The difference between AI becoming a domain expert or not may have less to do with how much information it has, and more to do with whether it knows what to pull out at the right moment. Which is, unfortunately, also the hardest part.

That also changes how I think about the value of enterprise data. The question may no longer stop at “How much data do we have?” What did that data mean in the situation where it was generated? What judgment did the expert make? And what did they do next?

Capturing that context may become the real advantage.

Giving AI domain expertise, then, may be less about feeding it ever more data and more about structuring the context and expert judgment that sit behind the work.

Next | What Happens When Context Enters the Physical World?

So what happens if we take the same three questions into manufacturing, semiconductors and defense?

The questions stay the same. The data used to answer them changes.

A factory has to understand vibration, temperature, and the current state of a Digital Twin. A semiconductor system has to understand the current design and the intent behind it. A defense system has to understand sensor inputs and the current mission state.

Domain Context starts moving beyond documents and databases into the live state of the physical world.

This time, instead of reading PDFs, the AI has to read machines, chips and battlefields.

In the next piece, I’ll take the same three questions into the world of Physical AI.

의료, 법률, 통신처럼 서로 다른 산업의 AI를 들여다보면 한 가지 질문이 생깁니다. AI에게 그 산업의 전문성을 가르친다는 것은 단순히 더 많은 산업 데이터를 학습시키는 일일까요? 이번 글에서는 MedGemma, Abridge, Harvey, CoCounsel, 그리고 제가 직접 참여했던 Telco LLM 사례까지 따라가며 산업 특화 AI가 실제로 무엇을 배우고 있는지 살펴봤습니다.

그 과정에서 반복해서 등장한 ‘현재 상황을 이해하고, 필요한 지식을 찾고, 전문가의 workflow를 따라가는’ 세 가지 질문을 통해 Domain Context가 어떻게 만들어지는지를 이해해보고자 합니다.

지난 글에서는 AI가 치르는 시험이 어떻게 바뀌고 있는지를 살펴봤습니다.

객관식 정답 찾기에서, 이제는 전문가의 장기 프로젝트를 처음부터 끝까지 완결할 수 있는지를 시험하는 단계까지.

수능시험에서 기업 입사 테스트, 혹은 승진 시험으로 바뀌고 있는 느낌이랄까요.

그런데 더 재미있었던 건 Agents’ Last Exam에서 SOTA 모델조차 통과율 0%였다는 점입니다. 가장 큰 문제는 Wrong StrategyDomain Knowledge Gap이었습니다.

결국 AI가 실제 전문가처럼 일하려면 그 산업을 이해하고, 주어진 상황에서 무엇이 중요한지 판단할 수 있어야 한다는 이야기였습니다.

그래서 오늘은 이 질문을 다뤄보려고 합니다.

AI에게 산업의 전문성은 어떻게 만들어줄 수 있을까?

먼저 산업을 나눠봤습니다

주요 산업을 하나씩 깊이 있게 들여다보려니, 어디서부터 어떻게 나눠볼지가 고민이었습니다. 그래서 크게 두 가지 카테고리로 나눠봤습니다.

  • Group A. 의료, 법률, 통신
  • Group B. 제조, 반도체, 국방

A는 방대한 전문 지식과 데이터를 기반으로 판단이 이루어지는 산업입니다. 동시에 이 정보들은 민감하고 보안과 규제의 영향을 크게 받습니다.

반면 B는 한 단계 더 물리 세계에 가깝습니다. AI가 공장의 기계 상태, 반도체의 현재 설계, 전장의 센서 데이터처럼 물리 시스템의 현재 상태를 이해하고 그 안에서 판단하고 행동까지 해야 합니다. 그야말로 최근 가장 주목받고 있는 Physical AI와 맞닿아 있는 영역이죠.

이번 2-1편에서는 먼저 의료, 법률, 통신을 보겠습니다. 각 산업의 핵심 플레이어들이 최근 어떤 모델과 제품을 내놓고 있는지 사례와 논문을 하나씩 따라가면서, 산업을 넘어 반복되는 공통점과 각 산업만의 차이가 무엇인지 찾아보려고 합니다.

그리고 다음 2-2편에서는 같은 질문을 제조, 반도체, 국방 산업으로 가져가 보겠습니다.

일단 의료부터 시작해보죠.

Healthcare | 의료

의학을 아는 것과 이 환자를 이해하는 것은 다르다

의료 AI를 전문화하는 가장 직관적인 방법은 AI에게 의학을 더 많이 가르치는 것입니다.

Google이 2026년 공개한 MedGemma 1.5가 대표적입니다. 일반 Gemma를 의료 텍스트와 이미지에 맞게 특화했고, CT와 MRI의 3D 영상, 병리 이미지, EHR 이해까지 범위를 넓혔습니다. 즉 범용 모델 자체를 더 좋은 ‘의료 모델’로 만드는 접근입니다.

MedGemma collection, 2D imaging부터 advanced imaging까지
의료 AI가 읽어야 할 정보도 넓어지고 있습니다. MedGemma 1.5는 텍스트와 2D 이미지를 넘어 CT, MRI, 병리 이미지, 환자의 과거 영상 변화까지 이해하는 방향으로 확장되고 있습니다.Source: Google Research, “Next-generation medical image interpretation with MedGemma 1.5” (2026)

그럼 의학책을 정말 많이 읽은 AI를 만들면 의사가 되는 걸까요?

2026년 의료 AI의 주요 업데이트들을 함께 놓고 보면, 답은 그렇게 단순하지 않아 보입니다.

OpenAI는 2026년 1월 ChatGPT for Healthcare를 출시했고, 7월에는 사용자가 의료 기록과 Apple Health를 연결할 수 있도록 Health in ChatGPT를 확장했습니다.

Health in ChatGPT 화면
“내 건강 상태를 알고 있는 AI”는 어떤 모습일까요? Health in ChatGPT는 의료 기록과 Apple Health를 연결해, 사용자의 질문을 실제 건강 정보의 맥락 안에서 이해할 수 있도록 하고 있습니다.Source: OpenAI, “Health in ChatGPT” (2026)

모델을 의료에 맞게 학습시키는 접근도 있지만, 주요 플레이어들은 동시에 모델 주변에 환자의 실제 맥락을 연결하기 시작했습니다.

서로 다른 제품이지만, 무엇을 해결하려는지 들여다보니 세 가지 질문으로 꽤 쉽게 정리할 수 있었습니다.

① 지금 무슨 일이 일어나고 있는가? 환자 상태

환자가 말합니다.

“요즘 숨이 너무 차요.”

범용 모델은 숨이 차는 원인을 수십 개 말할 수 있습니다. 하지만 의사에게 필요한 것은 가능한 질환 목록이 아니라 이 환자가 왜 숨이 찬지입니다. 나이와 병력, 복용 중인 약, 이전 진료 기록, 최근 검사 결과, 영상까지 함께 봐야 합니다.

같은 ‘숨이 차다’라는 말이라도 20대 운동선수와 70대 심장질환 환자에게는 전혀 다른 의미일 테니까요.

Abridge는 진료실에서 나누는 현재 대화뿐 아니라 응급실 기록, 간호 평가, 검사 결과, 영상, 이전 진료 기록을 연결해 진료 전에 환자의 상태를 구성합니다. OpenAI와 Anthropic도 각각 의료 기록과 건강 데이터를 범용 AI가 읽을 수 있는 현재의 개인 맥락으로 연결하고 있습니다.

즉 의료 산업에서 ‘지금 무슨 일이 일어나고 있는가?’에 답하려면 AI가 환자의 현 상태를 얼마나 제대로 구성할 수 있는지가 중요합니다. 여러 플레이어가 지금 이 환자 상태를 어떻게 데이터화하고, 필요한 순간에 모델이 읽을 수 있게 할지를 서로 다른 방식으로 풀고 있는 것이죠.

② 문제를 해결하기 위해 필요한 지식은 무엇인가? 임상 지식

환자의 상태를 안다고 끝나는 것도 아닙니다. 그 상태에서 문제를 해결하기 위해 어떤 의학 지식이 필요한지를 찾아야 합니다.

Abridge는 UpToDate에서 시작해 NEJM, JAMA, American Heart Association 등으로 근거 자료를 확대했고, 2026년에는 병원별 진료 지침과 치료 경로까지 반영하는 방향으로 확장했습니다.

OpenAI도 ChatGPT for Healthcare에서 동료심사 연구, 공중보건 지침, 임상 가이드라인을 출처와 함께 검색하고, 병원 자체의 승인된 치료 경로와 운영 지침까지 연결하도록 설계했습니다.

여기서 중요한 건 의학 지식을 ‘무작정 많이 넣는 것이 장땡’이 아니라는 겁니다. 핵심은 환자의 현재 상태를 먼저 이해하고, 그 위에 ‘지금 당장’ 문제를 해결하는 데 필요한 근거를 가져오는 것입니다.

의학 백과사전을 통째로 펼치는 게 아니라, 지금 이 환자에게 필요한 페이지를 정확히 펼치기 위한 시도인 셈이죠.

③ 전문가는 어떻게 처리하는가? 진료 워크플로

마지막은 실제 일하는 방식입니다. 의사는 정보를 읽고 답변만 생성하지 않습니다.

진료 전에 환자의 이력을 파악하고, 진료 중에는 증상과 근거를 바탕으로 판단하고, 이후에는 검사와 처방, 기록, 후속 관리로 이어갑니다.

Abridge가 2026년 제품을 진료 전 준비 → 진료 중 지원 → 진료 후 기록과 후속 업무까지 연결한 이유도 여기에 있습니다. OpenAI는 의뢰서, 사전승인, 환자 안내처럼 반복되는 임상 업무를 재사용 가능한 템플릿으로 만들고 있고, Anthropic도 환자 기록, 보험 기준, 임상 지침을 함께 확인하는 의료 업무용 스킬을 공개했습니다.

그래서 의료 AI를 따라가면서 처음 보인 구조는 이랬습니다.

환자 상태임상 지식진료 워크플로

이 환자에게 지금 무슨 일이 일어나고 있는지 파악하고, 문제를 해결하는 데 필요한 의학 지식을 가져오고, 실제 진료 흐름 안에서 사용할 수 있어야 했죠.

의료 산업에서 가장 애를 먹고 있는 건 바로 이 세 번째 질문에 답하는 것입니다. 의사들의 Tacit knowledge가 크게 작용하기 때문입니다. Abridge와 OpenAI 모두 최종 판단은 의료진에게 남겨두고 있습니다. AI가 정보를 잘 정리하는 것과 실제 의료 판단을 안정적으로 대신하는 것 사이에는 여전히 큰 간격이 있는 셈이죠.

Legal | 법률

리걸 산업에서도 묘하게 익숙한 구조가 다시 나타났습니다

다음으로 Harvey, Thomson Reuters, LexisNexis의 최근 업데이트를 봤습니다. 처음에는 의료와 전혀 다른 산업이라고 생각했습니다.

그런데 제품을 뜯어보니 질문이 거의 그대로 반복됐습니다.

① 지금 무슨 일이 일어나고 있는가? 사건 상태

의료에서 환자의 상태를 알아야 했다면, 법률에서는 현재 사건이나 거래의 상태를 알아야 합니다.

M&A 변호사에게 계약서를 한 장 주고 “우리 고객에게 위험한 부분을 찾아주세요”라고 해도 충분하지 않습니다. 어떤 거래인지, 고객이 누구인지, 어느 관할권인지, 이전 계약안에서는 무엇이 달랐는지, 협상 과정에서 어떤 입장을 주고받았는지를 함께 봐야 합니다.

Harvey는 그래서 문서 하나에 질문하는 수준을 넘어, 여러 자료와 조직의 문서관리시스템을 에이전트가 함께 읽을 수 있도록 연결하고 있습니다. 최근 Harvey Agents 역시 조직의 문서와 시스템에 접근해 복수의 자료를 하나의 업무 맥락 안에서 처리하도록 설계돼 있습니다.

Harvey Agents가 조직 문서를 오가며 실사 요청 목록을 처리하는 화면
Harvey Agents는 iManage 같은 조직의 문서 시스템과 여러 자료를 함께 연결해, 하나의 matter context 안에서 필요한 정보를 찾아 작업합니다.Source: Harvey, “Agents”

의료에서 환자 상태를 구성했다면 법률에서는 사건 상태를 구성하는 셈입니다.

환자가 사건으로 바뀌었을 뿐, 지금 정확히 무슨 일이 벌어지고 있는지를 파악하기 위한 사투는 산업 불문하고 계속되는 과제인가 봅니다.

② 문제를 해결하기 위해 필요한 지식은 무엇인가? 법률 및 로펌 노하우

두 번째 질문도 비슷합니다. 현재 사건을 해결하려면 어떤 법과 판례가 적용되는지 알아야 합니다.

Thomson Reuters의 Westlaw는 변호사들이 법령과 판례를 검색할 때 사용하는 대표적인 전문 법률 데이터베이스인데, 2026년 차세대 CoCounsel은 이런 검증된 자료를 AI와 직접 연결했습니다. 즉 인터넷에서 그럴듯한 답을 찾는 대신, 변호사가 실제로 사용하는 법률 자료에서 필요한 근거를 찾아오게 만드는 것이죠. LexisNexis도 자체 법률 데이터베이스를 AI와 연결하는 비슷한 방향으로 가고 있습니다.

CoCounsel의 검증 화면
법률에서는 그럴듯한 답보다 “무엇을 근거로 말했는가”가 중요합니다. CoCounsel은 AI가 제시한 주장과 실제 법률 근거가 맞는지 다시 확인할 수 있도록 검증 기능을 강화하고 있습니다.Source: Thomson Reuters, “CoCounsel Legal June 2026 releases”

그런데 법률에서는 여기에 한 층이 더 필요했습니다.

공개된 법률 지식 위에 로펌의 축적된 지식이 쌓인 2층 구조
법령과 규정, 판례는 모두가 딛고 서는 1층입니다. 로펌마다 조언이 달라지는 지점은 그 위에 얹힌 2층 — 과거 사건, 선호 문구, 내부 기준, 축적된 판단입니다.

비슷한 M&A 거래라도 고객이 감수하려는 위험, 협상에서 반드시 지키려는 조건, 과거에 수용했던 문구에 따라 로펌의 조언은 달라질 수 있습니다.

그래서 2026년 5월 Harvey와 DeepJudge는 법뿐 아니라 로펌의 과거 업무 결과물과 판단을 AI가 활용하도록 하는 Institutional Intelligence를 내세웠습니다.

Harvey와 DeepJudge의 Institutional Intelligence 파트너십 발표
법전은 모두에게 열려 있어도, 로펌의 경험은 그렇지 않습니다. Harvey와 DeepJudge는 과거 사건에서 축적된 조직의 판단과 노하우까지 현재 AI 업무에 연결하려 하고 있습니다.Source: Harvey, “DeepJudge and Harvey partner to power AI agents with Institutional Intelligence” (2026)

Westlaw 비밀번호를 받았다고 20년의 경험까지 같이 다운로드되는 건 아니니까요.

결국 법률 AI는 ‘법을 많이 아는 AI’를 넘어, 이 사건에 필요한 법을 찾고 우리 조직이라면 어떻게 판단할지까지 알아야 하는 셈입니다.

③ 전문가는 어떻게 처리하는가? 법률 워크플로

그리고 2026년 법률 AI에서 가장 뚜렷하게 보이는 변화는 ‘답변 생성’에서 ‘업무 수행’으로의 이동입니다.

Harvey는 3월 Agent Builder를 공개해 특정 팀이나 업무 분야가 자신들의 템플릿과 내부 절차를 재사용 가능한 에이전트로 만들도록 했습니다. 예를 들어 M&A 팀이라면 단순히 “이 계약서 검토해줘”라고 매번 요청하는 대신, 이런 흐름을 만드는 겁니다.

계약서 확인주요 위험 조항 추출기존 표준 문구와 비교수정안 작성검토가 필요한 쟁점 정리

LexisNexis 역시 Protégé에 미리 설계된 법률 워크플로와 Custom Workflow Builder를 추가해, 로펌이 자신의 업무 방식을 반복 가능한 AI 프로세스로 만들 수 있도록 하고 있습니다.

Protégé의 사전 설계 법률 workflow 목록
질문 하나에 답하는 AI에서, 일을 순서대로 수행하는 AI로. Protégé는 로펌의 실제 업무 방식과 기준을 반복 가능한 workflow로 옮기고 있습니다.Source: LexisNexis, “Global launch of Protégé AI workflows for legal professionals” (2026)

ChatGPT에게 계약서를 던지는 것과 실제 주니어 변호사에게 일을 맡기는 것의 차이를 줄이려는 시도라고 보면 조금 더 이해가 쉽습니다.

즉, 법률에서 도메인 특화를 위해 만들어가는 패턴은 이렇습니다.

사건 상태법률 및 조직 지식법률 워크플로

의료의 구조와 상당히 닮아 있습니다.

환자 상태임상 지식진료 워크플로

그 공통점이 보이시나요?

여기까지 보니 의료에서 나타난 패턴이 단순한 우연처럼 보이지 않기 시작했습니다.

Telecom | 통신

통신 특화 AI 역시 결국 이 세 가지 질문에 대한 답이라고 볼 수 있습니다

통신 산업은 제가 7년 반 동안 몸 담았던 산업입니다.

2024년 4대륙 통신사 간 AI 연합인 GTAA, Global Telco AI Alliance의 초기 런칭부터 Telco LLM 공동 설계와 GTM, 그리고 Syntelligence AI 설립으로 이어지는 과정에 직접 참여했습니다.

실제로 GTAA의 출발점도 한국의 SK텔레콤, 유럽의 Deutsche Telekom, 중동의 e&, 동남아의 Singtel, 일본의 SoftBank가 다국어 Telco LLM을 공동 개발해 각 통신사의 고객 서비스에 활용한다는 구상이었습니다.

아직도 5개 통신사가 거국적으로 한 테이블에 모였던 첫 회의의 긴장감을 잊지 못합니다.

모든 게 첫 시도다 보니 당시 분위기는 G20 정상회의급으로 비장했죠.

Global Telco AI Alliance 창립 서밋에 모인 다섯 통신사 경영진
2024년 MWC Barcelona. 서로 다른 시장과 언어를 가진 다섯 통신사가 하나의 Telco LLM을 만들기 위해 처음 한 테이블에 모였습니다.Source: SK Telecom Newsroom (2024)

당시에는 ‘통신을 잘 아는 LLM’을 만들면 꽤 많은 문제가 풀릴 거라고 생각했습니다.

하지만 실제 고객 서비스용 AI를 설계할수록 통신을 많이 아는 모델과 통신 일을 잘하는 모델은 다르다는 게 분명해졌습니다.

통신 특화 LLM을 개발하는 과정에서 다섯 개 통신사와 한 테이블에 앉아 끝장 토론을 이어가다 보니, 결국 모든 논의는 다시 같은 세 질문으로 수렴했습니다.

① 지금 무슨 일이 일어나고 있는가? 가입자 / 네트워크 상태

고객이 말합니다.

“집에서 와이파이 인터넷이 안 돼요.”

와이파이 문제라는 의도를 이해하는 것 자체는 어렵지 않습니다. 실제 업무는 여기서부터입니다.

같은 “와이파이가 안 돼요”라도 원인은 단말 설정일 수도 있고, 공유기일 수도 있고, 집으로 들어오는 회선이나 통신사 네트워크일 수도 있습니다.

즉, 통신 산업에서 이 ‘현재 상태’를 이해하기 위해서는 가입자 정보뿐 아니라 실제 네트워크와 장비의 실시간 상태까지 함께 봐야 하죠.

그래서 GTAA에서도 실시간 정보를 모델 안에 미리 학습시키는 것이 아니라, 문제가 발생한 순간 각 통신사의 시스템에서 필요한 최신 상태를 가져와 AI가 이해할 수 있도록 만드는 것이 중요했습니다.

문제는 다섯 통신사의 시스템이 모두 달랐다는 겁니다.

그래서 먼저 고객의 의도, 가입 상품과 회선 상태, 단말과 장비 정보, 현재 네트워크 상태처럼 AI가 판단에 필요한 핵심 정보를 공통으로 정의했습니다. 그리고 각 통신사의 서로 다른 가입자 시스템과 네트워크 운영 시스템에서 같은 의미의 데이터를 찾아 이 공통 스키마에 매핑했습니다.

서로 다른 통신사의 데이터를 AI가 같은 언어로 읽을 수 있도록 의미를 맞춘 셈입니다.

② 문제를 해결하기 위해 필요한 지식은 무엇인가? 통신 및 통신사 지식

가입자와 네트워크의 현재 상태를 공통된 형식으로 읽을 수 있게 만들고 나니, 바로 다음 질문이 나왔습니다.

“어디까지 공통으로 만들고, 어디부터 각 통신사에 맞춰야 할까?”

인터넷 장애, 요금 분쟁, 네트워크 문제 같은 공통 개념과 의도 체계는 공유할 수 있습니다. 하지만 실제 상품 구성, 사용하는 장비, 네트워크 환경, 정책 예외, 에스컬레이션 기준은 통신사마다 다릅니다. 그래서 공통적인 Telecom Intelligence 위에 각 통신사의 지식과 정책을 다시 얹는 구조가 필요했습니다.

2026년 TeleCom-Bench도 통신 기초 지식과 3GPP뿐 아니라 실제 장비와 상품에 관한 독점 지식, 실제 운영 workflow까지 함께 평가합니다.

TeleCom-Bench 프레임워크, Comprehension과 Application으로 구분
TeleCom-Bench는 ‘통신 지식을 이해하는 능력’과 ‘그 지식을 실제 업무에 적용하는 능력’을 별도의 영역으로 나눠 평가합니다.Source: TeleCom-Bench (2026)

단순히 “5G를 아는가?”에서 “지금 발생한 문제를 해결하기 위해 실제 통신 환경의 지식을 제대로 꺼내 사용할 수 있는가?”로 평가가 이동한 셈입니다.

교수님이 이번 시험은 오픈북이라고 했을 때 쾌재를 불렀지만, 결국 오픈북도 공부 안 하고 보면 죄다 틀린다는 걸 얼마 안 가 깨달았던 과거가 문득..

③ 전문가는 어떻게 처리하는가? 운영 워크플로

그리고 제가 가장 어렵다고 느꼈던 부분입니다.

5개사가 몇 날 밤을 새워 가며 가입자와 네트워크 상태를 공통된 형식으로 읽는 데이터 스키마에 합의했을 때 속으로 만세를 외쳤습니다.

오케이.. 이제 끝났다. 진짜 다 왔다..

하지만 진정한 챌린지는 그 다음 질문이었죠.

잠깐만, 근데 판단 로직은 어떻게 하지? (정적)

통신 전문가 머릿속에는 늘 ‘판단 로직’이 있습니다.

  • “다른 기기도 안 되면 단말보다 공유기나 회선을 먼저 본다.”
  • “공유기는 정상인데 회선 상태가 비정상이면 네트워크 쪽을 확인한다.”
  • “원격 조치로 해결되지 않으면 다음 단계로 에스컬레이션한다.”

TeleCom-Bench에서 현재 모델들은 의도 파악이나 주요 정보 추출에서는 90%를 넘는 성능에 도달하지만, 실제 해결책을 만들어내는 단계에서는 최고 모델조차 약 30% 수준에 그쳤습니다. 연구진은 이를 Execution Wall이라고 불렀습니다.

TeleCom-Bench 결과 표, Solution Generation 점수가 크게 낮음
문제를 이해하는 것과 실제로 해결하는 것은 달랐습니다. TeleCom-Bench에서 모델들은 의도 파악에는 90% 이상의 성능을 보였지만, 실제 해결책 생성에서는 약 30% 수준까지 성능이 떨어졌습니다.Source: TeleCom-Bench (2026)

결국 저희도 각 질문에 대한 답을 단계별로 고민하고 있었던 겁니다.

가입자 / 네트워크 상태통신 및 통신사 지식운영 워크플로

그리고 GTAA에서 시작된 합작법인이 2026년 Syntelligence AI로 공식화되면서 이러한 변화를 녹여냈습니다. 현재 Syntelligence는 네트워크 인텔리전스, 통신 신뢰도, 고객 경험처럼 실제 운영 문제를 해결하는 production-ready AI를 핵심 목표로 내세우고 있습니다.

세 번째 반복되는 것을 보니, 이제 우연 같지는 않습니다

쓰는 데이터도 다르고, 산업의 규칙도 완전히 다르지만 결국 세 산업이 같은 질문에 답하고 있었던 셈입니다.

① 지금 무슨 일이 일어나고 있는가?

  • 의료: 환자 상태
  • 법률: 사건 상태
  • 통신: 가입자 / 네트워크 상태

② 문제를 해결하기 위해 필요한 지식은 무엇인가?

  • 의료: 임상 지식
  • 법률: 법률 및 조직 지식
  • 통신: 통신 및 통신사 지식

③ 전문가는 어떻게 처리하는가?

  • 의료: 진료 워크플로
  • 법률: 법률 워크플로
  • 통신: 운영 워크플로

산업은 계속 바뀌는데, 고민하는 질문들은 사실상 똑같네?

이제 패턴이 보이시나요?

의료: 환자 상태임상 지식진료 워크플로

법률: 사건 상태법률 및 조직 지식법률 워크플로

통신: 가입자 / 네트워크 상태통신 및 통신사 지식운영 워크플로

저는 이 세 질문에 대한 답이 산업별로 공통적으로 반복된다는 점에서, 이것이 결국 ‘도메인 맥락 레이어’, Domain Context Layer의 핵심 구조를 이룬다고 생각해 보게 됐습니다.

의료, 법률, 통신을 가로지르는 Domain Context Layer
산업은 달라도 구조는 반복됐습니다. 의료에서는 환자, 법률에서는 사건, 통신에서는 가입자와 네트워크를 이해한 뒤 필요한 지식을 찾고, 마지막으로 전문가의 workflow를 따라갑니다.

My Read | 산업 특화 AI의 경쟁은 결국 ‘맥락을 누가 더 잘 움켜쥐느냐’의 싸움이 될 것입니다

처음에는 산업 특화 AI를 이렇게 단순하게 생각할 수 있습니다. 의료 데이터를 넣으면 의료 AI, 법률 데이터를 넣으면 법률 AI, 통신 데이터를 넣으면 통신 AI.

그런데 실제 제품들을 따라가 보니 그렇게 단순하지 않습니다. 도메인 맥락 레이어를 만들기 위해서는 세 질문에 대한 답이 반드시 필요하고, 그 답을 만드는 재료는 산업마다 완전히 다릅니다.

의료는 환자의 기록과 검사, 임상 근거를 사용합니다. 법률은 사건 자료와 신뢰할 수 있는 법률 정보, 로펌의 과거 판단을 사용합니다. 통신은 가입자와 네트워크 상태, 통신사별 상품과 정책을 사용합니다.

각자 다른 데이터를 쓰고, 다른 시스템과 규제를 가지고, 다른 방식으로 일합니다. 그런데 AI에게 전문성을 만들어주는 방식은 놀랄 만큼 비슷했습니다.

산업 전체를 모델 안에 집어넣는 것보다, 전문가가 실제 판단할 때 사용하는 맥락을 AI 주변에 얼마나 잘 재구성할 수 있느냐가 더 중요합니다.

생각해보면 전문가도 모든 것을 동시에 떠올리지는 않습니다. 의사는 이 환자에게 필요한 근거를 보고, 변호사는 이 사건에 필요한 법과 로펌의 과거 판단을 가져오고, 통신 전문가는 이 가입자와 네트워크 상태에 맞는 정책과 해결 경로를 선택합니다.

AI가 도메인 전문가가 될 수 있냐 없냐의 차이는 정보의 양보다, 지금 무엇을 꺼내 써야 하는지를 아는 데 있는지도 모르겠습니다. 근데 그게 제일 어려운데 쩝

그래서 기업 데이터의 가치도 단순히 ‘얼마나 많이 가지고 있는가’에서 끝나지 않을 것 같습니다. 그 데이터가 어떤 상황에서 무엇을 의미했고, 그때 전문가는 어떤 판단을 내리고 다음에 무엇을 했는가. 이 맥락 데이터를 쥐는 것이 이제는 핵심입니다.

즉, AI에게 산업의 전문성을 준다는 것은 더 많은 데이터를 넣는 문제를 넘어, 이런 맥락과 전문가의 판단 방식을 얼마나 잘 구조화할 수 있느냐의 문제일지도 모릅니다.

Next | 맥락이 물리 세계로 나오면?

그렇다면 같은 세 질문을 제조, 반도체, 국방에 가져가면 어떻게 될까요?

공장에서는 기계의 진동과 온도, Digital Twin 속 현재 상태를 읽어야 하고, 반도체에서는 현재 설계와 설계 의도를, 국방에서는 센서와 임무 상태를 이해해야 합니다.

즉 Domain Context가 문서와 데이터베이스를 넘어 물리 세계의 실시간 상태로 확장되기 시작합니다.

이번에는 PDF 대신 기계와 칩, 그리고 전장을 읽어야 하는 셈입니다.

다음 편에서는 같은 세 질문을 들고 Physical AI의 세계로 가보겠습니다.