The Invisible Cost of Disorder: How Information Architecture Prevents SEO Losses

by | Mar 16, 2026 | Knowledge Organization | 0 comments

Esse artigo pode ser lido em: Portuguese (Brazil)

In this article, I’m going to write about a topic that, at first glance, seems far removed from SEO. We usually talk about technical topics, indexing, algorithm updates and, more recently, AI. But it’s another kind of IA that I want to address here: Information Architecture. But let’s look at it from a new point of view: come with me!

In corporate ecosystems of greater complexity and in business environments that are increasingly digitized, information architecture should not be seen as a simple organizational layer in interface development. In fact, it is a strategic powerhouse against losses that exceed the millions and are generated by the continuous loss of productivity.

This phenomenon occurs when employees of organizations cannot retrieve information that is vital to their work, or when customers cannot find answers to their queries and searches in the search tools on the websites of those organizations.

User Experience, Information Architecture, and Semantic SEO specialists operate at the exact intersection between human cognition and the increasingly intricate Web data infrastructure. It is in this scenario that categorization must act as a primary reducer of cognitive load and, simultaneously, as the great engine that drives search systems toward semantics—whether the internal search of a portal or the indexing carried out by today’s search engines.

When data is not organized in a logical structure, Natural Language Processing suffers, algorithms fail to understand the meaning of the content, and the business’s organic visibility goes from bad to worse.

Fundamentals of Categorization: The Science of Organizing Objects and Entities

Categorization and information architecture are the strategies you’re looking for to increase digital findability and avoid the invisible cost of disorder.

Want to read more about strategies and tactics in SEO?

Introduction to Classification Logic in the Digital Context

Categorization is the pillar that underpins human cognition, the foundation that allows our brain to process massive volumes of information by grouping entities by similarity and meticulously distinguishing their dissimilarities. Inside your skull you have the best categorization machine ever invented.

In digital environments and in Information Science itself, this organizational logic is what separates an intuitive and enriching user journey from absolute informational chaos.

For the information architect and the SEO professional, organizing information means mapping the user’s mental model, in order to reduce the effort of choice, transforming raw data into structured, quickly retrievable assets, effectively and efficiently.

When we deal with modern search algorithms, such as BERT, the machine needs to understand which “entity” a piece of content belongs to in order to deliver it as the best answer to a query. Without an efficient classification logic, the content produced becomes invisible and loses its value.

Analysis of Concept Types and Attributes

Have you heard of the NISO Z39.19 guidelines?

The ANSI/NISO Z39.19-2005 (R2010) guidelines establish essential standards for the construction, formatting, and management of monolingual controlled vocabularies, including thesauri, lists, synonym rings, and taxonomies. The focus of the guidelines is the consistent representation of content objects to facilitate information retrieval in knowledge systems.

A curiosity: did you know you can use NISO Z39.19-2005 as a basis for building the new darlings of AI tools, ontologies? Access this article at cip.brapci.inf.br/download/135118 and read how to do it.

Returning to our conversation about information organization: we know that structuring a competent and optimized database requires the correct identification of attributes and classes, which enables the implementation of multidimensional faceted search, a vital functionality for extensive catalogs like the ones we’ve seen in e-commerce.

Based on NISO’s (National Information Standards Organization) standardized guidelines, I summarized the seven essential types of concepts that we, who work with information representation, need to know in order to structure any taxonomy:

  • Things: refer to physical objects, tangible entities, and their constituent parts. In e-commerce, it can be a “notebook” or a “processor.”
  • Materials: substances from which things are formed. For example, specifications such as “aluminum,” “glass,” or “silicon.”
  • Activities: processes, actions, or operations performed. In the web environment, they represent interactions such as “buy,” “review,” “compare,” and “share.”
  • Events: occurrences or phenomena situated in time, such as “Black Friday,” “SEO Course,” or “Campaign Launch.”
  • Properties: characteristics, states, or qualities inherent to an object. It can be size, primary color, exact weight, or storage capacity.
  • Disciplines: areas of study or broad branches of knowledge. Here, comprehensive thematic categories enter, such as “Library Science,” “Software Engineering,” and “Digital Marketing.”
  • Measures: units of dimension, scale, or quantity, such as “centimeters,” “gigabytes,” “kilometers,” or financial currencies.

Notice that, with these seven categories in hand, you can already organize all the information in a product catalog. Work together with your systems or software development team and you’ll be able to create a cutting-edge search or product recommendation system.

Similarity and Dissimilarity Criteria: The Strategic Impact on Data Retrieval

Before moving forward, I need to address these two concepts from the point of view of Information Science. We need to understand that the concepts of similarity and dissimilarity are part of the foundations of organization, retrieval, and representation of information. They are not merely subjective perceptions, but a practical way that allows systems (human or artificial) to identify relationships between documents, terms, or entities.

So I present a technical and reflective definition of these two concepts:

Similarity

Similarity is the degree of correspondence, proximity, or affinity between two informational objects. In Information Science, it is often treated from two perspectives:

  • Structural similarity: focuses on the form or physical occurrence of elements (e.g., two articles that share the same keywords).
  • Semantic similarity: focuses on meaning. It occurs when two terms or documents address the same concept, even if they use different languages or terms (synonymy).

Mathematically, similarity is frequently calculated in a vector space (does it remind you of how AI models work?), in which documents are represented by vectors. The most common metric is cosine similarity, which measures the angle between two vectors:

sim(A,B)=cos(θ)=ABAB\text{sim}(A, B) = \cos(\theta) = \frac{A \cdot B}{\|A\| \|B\|}

The closer to 1, the greater the similarity between objects.

Dissimilarity

Dissimilarity is the measure of distance, difference, or divergence between objects. In practice, it is the inverse of similarity, but it has a fundamental strategic value in categorization and classification.

While similarity groups, dissimilarity separates, and is essential for:

  • Avoiding redundancy: in search systems, showing very similar results can be inefficient; dissimilarity helps ensure diversity of results.
  • Outlier identification: detecting information that does not fit any established pattern.

In metric terms, dissimilarity is often expressed as a “distance.” Euclidean distance is one of the ways to calculate this divergence:

d(x,y)=i=1n(xiyi)2d(x, y) = \sqrt{\sum_{i=1}^{n} (x_i – y_i)^2}

The dialectic between similarity and dissimilarity enables the creation of taxonomies and ontologies.

  1. Clustering: objects with high internal similarity and high external dissimilarity form a solid class or category.
  2. Cognitive load: efficient information organization uses these concepts to reduce the user’s mental effort. When the similarity between menu options is very high (ambiguity), cognitive load increases, as the user cannot distinguish the correct route.
  3. Information retrieval: modern search engines use natural language processing (NLP) and large-scale language models to refine this perception, going beyond simple word counting to understand “conceptual proximity.”

Point for reflection: in Information Science, nothing is “equal,” only “highly similar.” Absolute identity is rare; we always work with degrees of approximation that define the relevance of an answer to a query.

Forming Groups by Similarity or Dissimilarity

The formation of logical, semantic, and functional groups depends strictly on the distinction between intrinsic characteristics (what the object or entity actually is in its ontology) and extrinsic characteristics (how it is used, perceived, or applied by the end user). The failure to define and isolate these criteria generates immense algorithmic “noise,” harming crawlers and severely degrading search precision.

But let us clarify these complicated concepts.

I often say that an e-commerce without a clear distinction between what a product is and what a product represents is merely a digital warehouse, not a sales strategy.

Let’s use an example of a high-demand insulated bottle (like a Stanley or similar). Imagine we are organizing this store’s taxonomy and semantics:

The Intrinsic View (The Ontology of the Object)

Here, the group is formed by the similarity of what the object actually is. It does not matter who buys it or for what purpose.

  • Characteristics: stainless steel, vacuum insulation, 500 ml capacity, screw-on lid.
  • Logical grouping: kitchen > containers > thermal bottles.

The Extrinsic View (User Perception and Application)

Here, physical dissimilarity is ignored in favor of functional similarity. The group is formed by context.

  • Scenario A (the camping enthusiast): the bottle is grouped with tents, sleeping bags, and lanterns. It does not “look like” a tent, but it serves the same extrinsic purpose: survival and outdoor comfort.
  • Scenario B (the office professional): the bottle is grouped with planners, desk organizers, and ergonomic mice. Here, it is a productivity and status accessory.

In my professional experience, I have noticed that the common mistake is trying to force the user to think only in ontology (intrinsic). If your customer wants “gifts for adventurous fathers,” they do not want to navigate through “stainless steel > 500 ml.”

If information architecture does not reflect this user cognitive load—who searches by use and not by raw material—the internal search system fails, NLP can’t connect the dots, and conversion plummets. Extrinsic similarity is what generates desire; intrinsic similarity is what validates the technical purchase.

So, strategically, if the categorized attributes are ambiguous, the query to the database or to a search engine will return out-of-context results, forcing the visitor into manual and extremely exhausting filtering, which, invariably, increases the bounce rate.

The rigorous identification of those NISO attributes is one of the factors that allows a faceted navigation system to perfectly differentiate “Activities” from “Disciplines” in the side navigation filters.

The Role of Entity Extraction in Content

Another indispensable point in these foundations is “Entity Extraction.” In Semantic SEO, we need to constantly identify entities (people, places, organizations, concepts) present in the full text of a document and ensure that they align with the concepts identified in the knowledge-domain analysis and with the taxonomy created for the site.

By applying Natural Language Processing routines, we make precise inferences about these entities, consolidating the semantic domain of the page and attesting to the search engine that our content is a solid authority in that field of knowledge. Semantic SEO was already using AI in SEO long before this topic became fashionable.


This is the first part of the article that addresses how information architecture is important for SEO projects. The second part will be published soon.


Alexander Rodrigues Silva

Alexander Rodrigues Silva

SEO Specialist and Author of the Book Semantic SEO

Hello, I am Alexander Rodrigues Silva, an SEO specialist and author of the book “Semantic SEO: Semantic Workflow.” I have been working in the digital universe for over two decades, focusing on website optimization since 2009. My choices have led me to delve into the intersection between user experience and content marketing strategies, always with a focus on increasing organic traffic in the long term. My research and specialization concentrate on Semantic SEO, where I investigate and apply semantics and connected data in website optimization. It is a fascinating field that allows me to combine my background in advertising with library science.

Blog Semântico
Privacy Overview

This website uses cookies so that we can provide you with the best user experience possible. Cookie information is stored in your browser and performs functions such as recognising you when you return to our website and helping our team to understand which sections of the website you find most interesting and useful.