Large Language Models become much easier to understand when we start with what they actually process: language.
Much of the confusion surrounding LLMs starts by asking whether they think, reason, understand, or possess knowledge.
Start with language instead.
Language is the systematic use of signs, syntax, and semantics to encode and decode information.
From that perspective, an LLM is best understood as a sophisticated Language Emulation Machine — or, more succinctly, a Langulator.
So what happens to signs, syntax, and semantics inside an LLM?
Signs, Syntax, and Semantics
The basic division is straightforward:
Signs provide denotation.
Words, symbols, sounds, and identifiers stand for or refer to things.
Syntax provides grammatical arrangement.
It determines how signs can be combined into well-formed expressions.
Semantics provides meaning and role implications.
It tells us what those signs and arrangements mean.
And pragmatics adds context — who is communicating, about what, when, where, and under what circumstances.
Consider:
I deposited the check at the bank.
versus:
We sat on the bank of the river.
Same sign: bank.
Different meaning.
Context makes the difference.
What Does an LLM Do?
An LLM doesn’t explicitly implement these language components the way a symbolic system does.
It learns their patterns from examples.
At a high level:
Signs → Tokens and numerical representations
Syntax → Learned structural regularities
Semantics → Learned contextual associations
Pragmatics → Runtime context
During training, the model encounters enormous numbers of examples of signs being arranged syntactically and used in contexts carrying semantic implications.
Those regularities become encoded numerically in the model’s parameters.
That’s why an LLM can produce grammatical sentences without consulting a grammar book.
It’s why it can distinguish a financial bank from a river bank without consulting an explicit ontology.
And it’s why reducing an LLM to “just next-token prediction” misses something important.
The prediction mechanism may be simple to describe.
What has been compressed into the function performing that prediction is not.
Implicit Semantics
Consider:
Paris is the capital of France.
An LLM has encountered countless linguistic patterns connecting:
Paris
France
capital
city
country
Europe
Those relationships become embedded in its learned numerical structure.
The model can therefore generate language that behaves as though it knows:
Paris → capital of → France
But that’s not necessarily stored as an explicit assertion.
Compare that with RDF:
<Paris> rdf:type <City> .
<Paris> <capitalOf> <France> .
<France> rdf:type <Country> .
Now the relationships are explicit.
They are identified, declared, machine-computable, and available for deterministic inference.
That’s the critical distinction.
LLMs learn implicit semantics from language.
Semantic Web technologies express explicit semantics as data.
Language Is the Map
This also helps explain why LLMs appear to “know” so much.
Human knowledge has been encoded in language for thousands of years.
Books, papers, specifications, documentation, source code, conversations, and Web pages are all linguistic artifacts carrying information.
Train sufficiently powerful mathematical machinery on enough of those artifacts and it learns patterns reflecting the information encoded within them.
In simple terms:
Human Knowledge
↓
Encoded as Language
↓
Training Corpus
↓
Statistical Learning
↓
LLM
↓
Language Emulation
Language is the map.
The things language describes are the territory.
An LLM becomes extraordinarily capable at navigating and generating the map.
That doesn’t make the map the territory.
Enter The Semantic Web Project
This is where LLMs and a Semantic Web fit together remarkably well.
An LLM gives us:
probabilistic language processing.
RDF and Linked Data give us:
explicit standardized identifiers and relationships.
Ontologies give us:
machine-computable semantics.
Logic gives us:
deterministic inference.
The division of labor becomes clear:
Human
↓
Natural Language
↓
LLM / Langulator
↓
RDF + Linked Data + Ontologies
↓
Data Spaces
↓
Deterministic Results
↓
LLM / Langulator
↓
Natural Language
↓
Human
The LLM handles the fuzzy language.
A Semantic Web handles explicit meaning and relationships.
The surrounding Agent harness determines what actions may actually occur.
Why This Matters for AI Agents
An AI Agent doesn’t need its LLM to memorize every fact, relationship, policy, or procedure.
The LLM can interpret natural-language intent.
The Agent harness can select the appropriate Skills and Tools.
Those Tools can operate against Data Spaces — databases, knowledge bases, filesystems, and APIs.
And a Semantic Web can provide the explicit, machine-computable Context Layer connecting everything.
That produces a much cleaner architecture:
LLM
Language
Probabilistic
+
Semantic Web
Context
Deterministic
+
Agent Harness
Action
Controlled
Each component does the job for which it is best suited.
The Missing Piece Wasn’t Intelligence
For decades, explicit semantic systems had a UI/UX problem.
Humans were expected to interact with identifiers, schemas, ontologies, query languages, and graph structures.
LLMs change that.
They put natural language at the top of the computing UI/UX stack.
Humans can speak naturally.
LLMs can handle the fuzziness of that language.
Semantic infrastructure can provide precise, verifiable context underneath.
And LLMs can translate the results back into language.
That’s the symbiosis.
LLMs don’t make explicit semantics obsolete. They make explicit semantics easier to use.
And perhaps that’s the simplest way to understand what an LLM really is:
Signs become tokens. Syntax becomes learned structure. Semantics becomes learned association. Context shapes interpretation.
That’s a remarkably powerful Language Emulation Machine.
But when we need persistent identity, explicit relationships, verifiable meaning, and deterministic inference, we still need something else.
Language for the fuzzy parts.
Semantic Web for the precise parts.
Loose coupling between them.
Related
- Charles Sanders Peirce — Theory of Signs — Signs, objects, interpretants, reference, and denotation.
- Theories of Meaning — Stanford Encyclopedia of Philosophy — Meaning, reference, compositionality, and the relationship between syntax and semantics.
- Claude Shannon — A Mathematical Theory of Communication — Encoding, transmission, and decoding of information, while explicitly distinguishing information theory from semantic meaning.
- Distributional Hypothesis — The foundation for deriving semantic relationships from patterns of linguistic context.
- Mikolov et al. — Distributed Representations of Words and Phrases and their Compositionality — Numerical vector representations that capture syntactic and semantic relationships.
- Vaswani et al. — Attention Is All You Need — The Transformer architecture underpinning modern LLMs and their context-sensitive processing of language.
- Devlin et al. — BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding — Context-dependent representations demonstrating how the meaning of a sign can be conditioned by surrounding signs.
- RDF 1.1 Semantics — W3C — The contrasting explicit-semantic model: IRIs, interpretations, formally defined relationships, and machine-computable entailment.
