What Information Science teaches modern SEO

Alexander Rodrigues Silva
Alexander Rodrigues Silva verified
Arquiteto de Informação e Engenheiro de Engenharia de Busca Semântica
calendar_today 26 de August, 2026 • timer 9 min de leitura • Revisão por Pares: Sim
What Information Science teaches modern SEO

Understanding the evolution of the semantic web requires tracking the transition from the chaos of unstructured data to the ontological precision of entities. This profoundly deep thought came to me, out of the blue, in the corridors of the university where I study, UFRGS.

I found myself reflecting on the evolution of our profession. After two decades, and a bit, working in advertising, developing digital ecosystems, and producing textual content for the web, I watched the internet transform from a chaotic repository of links into a structure that seeks to mimic the way we think. My mind as a specialist focused on SEO and information architecture realized that our discipline has changed irreversibly, and I found myself grateful for having had contact with Information Science.

The digital library and the need for ontological order

In Information Science studies, one learns early on that the field is dedicated to the complete information cycle:

Collection, classification, storage, and methodical retrieval.

Historically, however, the web was born devoid of this structure. Organizing the internet in the early 2000s was similar to the task of managing a monumental collection without a bibliographic catalog or decimal classification system. Information retrieval depended strictly on exact matches, and we would type a sequence of characters hoping the algorithm would find a document with that exact syntactic combination, without consideration for context or disambiguation. It was an ocean of raw data.

The great shift began with the emergence and application of algorithms, but the rupture of this model occurred with the advancement of Generative Artificial Intelligence and the Transformer architecture, which enabled the emergence of large-scale language models, the so-called LLMs. These models allowed the transition from a purely syntactic flow to an organization with pinpoint semantic precision.

While legacy systems foundered in the noise of informational superabundance, modern algorithms began using semantic networks and complex taxonomic hierarchies to process information. The web could cease being an agglomeration of isolated documents and become an interconnected knowledge graph.

In this scenario, the transition from a focus on keywords to structured concepts became indispensable for brand survival in the era of semantic search. Investment in advanced language models involves colossal sums; faced with this level of sophistication, digital authority no longer belongs to those who simply replicate terms in a text, but to those who establish real semantic interoperability. For search engines, the shift from “strings to things” (that is, to entities) is the dividing line that separates disposable content from Authority.

The linguistic sign and the raw material of search intent

All this dynamic converges on the foundational unit of language: the linguistic sign.

Deciphering the user’s real search intent requires mastery of semantics. As information architects, we understand that each search query is a symbolic representation seeking validation. The linguistic sign, according to the Saussurean tradition, is the binary unit composed of the inseparable relationship between the signifier (the physical form or digital characters of the query) and the signified (the mental concept evoked), and no algorithm or model can break this relationship.

To illustrate the limitation of systems based solely on syntax, we frequently turn to John Searle’s “Chinese Room” experiment, which demonstrates how the manipulation of syntactic symbols does not guarantee real semantic understanding.

What is the Chinese Room experiment?

The “Chinese Room” thought experiment, proposed by John Searle in 1980, perfectly illustrates a decisive debate about the true understanding of machines. He argues against the idea that a computer can truly think or develop a conscious mind.

To understand the mechanics of the idea, imagine you are locked in a room and do not know how to speak or read a single word in Chinese. Inside this isolated space, you receive a large instruction book in Portuguese that methodically details how to combine Chinese symbols, and this manual dictates exact processing rules, such as: “if you receive symbol x, write and return symbol y.”

The interaction occurs when people outside the room begin sliding notes with questions in Chinese under the door. Even without understanding the language itself, you follow the manual very efficiently, find the correct syntactic combination, and return the answer. For those outside, the illusion is undetectable: it seems the person inside is fluent in Chinese and conversing naturally. But the reality, however, is that you continue to understand absolutely nothing of the meaning of those symbols, operating only mechanically.

The significance of this argument touches on an essential point I always try to relate to my research in Library Science: the enormous difference between syntax and semantics. Searle demonstrates that computers operate strictly at the level of syntax, focusing on the formal manipulation of rules and symbols based solely on their structural forms. These machines do not possess semantics, which means they lack real understanding or intentionality of what these symbols represent for the world and for human experience.

This is a scathing critique of the so-called strong artificial intelligence view, which argues that the human mind operates merely as a program and that complex software would inevitably result in a conscious machine. For Searle, simulating intelligence does not equate to possessing intelligence.

AI is indeed a field dedicated to seeking methods that multiply the capacity to solve problems, but any current large-scale language model, however impressive and persuasive it may seem, still acts like the occupant of that closed room. It is processing symbols brilliantly, but groping in the dark when it comes to real meaning.

Modern LLMs attempt to overcome this limitation by mimicking the relationship of the sign through billions of mathematical parameters, seeking what the literature calls grounding—the anchoring of the sign in a verifiable and factual context, as I mentioned previously—and the absence of solid grounding results in so-called hallucinations.

In semantic SEO, failing to align the signifier with the correct signified leads to serious indexing and disambiguation errors.

This anchoring dynamic is aligned with the principle of predictability in information and learning systems. When structuring an informational environment, clarity and consistent reinforcement are fundamental. Just as animal training based on clear routines creates predictable connections, providing hierarchical data markup to search crawlers eliminates uncertainties. When the crawler analyzes a web page, it seeks explicit declarations of identity and context to validate the entities represented.

The term as conceptual specialization of the lexicon

In the specialized search ecosystem, one observes the metamorphosis of the word into term.

While the word belongs to the general lexicon and carries polysemic sampling, the term is the unit of precise signification that designates a specific concept within a well-delimited semantic domain. The term neutralizes the ambiguity of language. For example, the word “transformer” in common usage refers to electrical devices or popular culture, whereas in the artificial intelligence domain the term “Transformer” rigorously designates a neural network architecture based on attention mechanisms.

With advances in Natural Language Processing, algorithms began treating terms as vectors in a multidimensional space. Conceptual proximity is measured by the vector distance between terms in the network. This approach expands the scope of informational optimization, shifting the focus from simple keyword repetition to the mapping and construction of dense and coherent semantic networks.

In the practice of the data analysis laboratory, the extraction and validation of these vector networks are performed through automation scripts that map semantic relevance and conceptual density. This process ensures that the selected terms provide the necessary ontological support to consolidate a document’s authority before retrieval algorithms.

The descriptor and the authority of controlled vocabulary

The apex of maturity in knowledge organization is the formal adoption of the descriptor. If the term specializes the word, the descriptor acts as the univocal element of a controlled vocabulary, standardized to represent concepts without ambiguity. In the semantic web, this function is performed primarily by structured data and structuring tools such as taxonomies.

The descriptor functions as the canonical key that connects the document to the knowledge graph (for example, the Google Knowledge Graph). Treating structured data merely as visual resources for the results page limits their potential; they are, essentially, the direct interoperability bridge between the site’s information architecture and global knowledge systems.

This formal validation is equivalent to institutional processes of registration and authentication. Just as an official document attests to legal relationships without room for ambiguous interpretations, markup with structured data formally declares the ontological relationships between a site’s entities (author, organization, topics, and products). It is the formalization of trust on the network.

The strategic parallel between Information Science and SEO

The convergence between Library Science, Information Science, and site optimization establishes the foundations of contemporary digital visibility. When designing the taxonomy of a digital ecosystem, one constructs an architecture capable of organizing knowledge into well-delimited structures, facilitating crawling, indexing, and understanding of the informational value offered.

When the search engine recognizes that a project respects the hierarchical and associative relationships between topics, progressing from the generic to the granular, greater reliability is attributed to the domain. Content ceases to be a set of disconnected texts to form a structured knowledge base.

This rigorous theoretical-practical structure guides the translation of concepts from Computer Science and Information Science for application in the market, connecting academic theory to the demands of managers and information architects.

The continuous sophistication of language models requires the transition from operational executors to information architects. Although technology provides unprecedented computational capacity, human curation and critical analysis of the intent underlying the data remain central differentiators.

In semantic search, sustainable success requires the provision of structured data that is clean and contextualized. Algorithms operate on mathematical representations, but depend on human analytical rigor for the correct modeling of knowledge.

The central question for structuring any digital project remains: in what way will the adopted taxonomic architecture reflect the necessary transition from ambiguous lexicon to the ontological consolidation of entities?

school FTS TRAINING PROGRAM

Semantic SEO Course

Master the Semantic Workflow (FTS) and position your content at the forefront of digital optimization in the era of Artificial Intelligence.

check_circle 12 Theoretical and hands-on lab modules
check_circle Access to the Graph Studio ecosystem
check_circle Bi-weekly direct mentoring with the author
Enroll in the Course arrow_forward
Alexander Rodrigues Silva
Sobre o Autor

Alexander Rodrigues Silva

Pesquisador em Ciência da Informação, especialista em SEO Semântico e autor do livro "SEO Semântico: Fluxo de Trabalho Semântico". Lidera as pesquisas da Semântico SEO em Porto Alegre focadas na formalização de dados e arquiteturas GraphRAG.

DIÁLOGO CONCEITUAL

Deixe sua Reflexão ou Pergunta

Seu endereço de e-mail não será publicado. Campos obrigatórios são marcados com *.

Blog Semântico

Privacy Overview

This website uses cookies so that we can provide you with the best user experience possible. Cookie information is stored in your browser and performs functions such as recognising you when you return to our website and helping our team to understand which sections of the website you find most interesting and useful.