Retrieval-Augmented Generation (RAG) systems are only as good as the data they ingest. While processing plain text documents is relatively straightforward, handling unstructured or semi-structured documents—specifically multi-column reports and complex tables—remains a significant bottleneck. When text flows laterally across columns or when tabular data loses its structural context during parsing, the resulting embeddings become noisy, leading to hallucinations and irrelevant retrieval results.
The Challenge of Non-Linear Text Structures
Standard parsers often read documents in a linear fashion, column by column or line by line. In a multi-column layout, this approach results in fragmented sentences. For example, a sentence might start at the bottom of the first column and continue at the top of the second. If not handled correctly, the chunking engine may split this sentence in half, destroying its semantic meaning. Furthermore, tables in PDFs or HTML often lack explicit row and column headers that repeat on every page, making it difficult for an LLM to understand the context of the data without significant preprocessing.
Strategy 1: Semantic Chunking with Contextual Overlap
To address fragmented sentences, we must move beyond simple character-based chunking. Semantic chunking involves splitting text based on natural breaks, such as paragraphs or complete sentences. When implementing this, it is crucial to introduce an overlap window. This ensures that even if a sentence is split, subsequent chunks retain the context from the previous one.
Here is a practical example using Python and LangChain to implement semantic chunking with overlap:
from langchain.text_splitter import RecursiveCharacterTextSplitter
# Define a splitter that respects sentence boundaries
text_splitter = RecursiveCharacterTextSplitter(
chunk_size=500,
chunk_overlap=50,
length_function=len,
separators=["\n\n", "\n", ". ", " ", ""],
)
# Input text from a multi-column report
report_text = """
The Q3 financial results indicate a 15% increase in revenue.
However, operating expenses have also risen due to
supply chain disruptions.
Net profit margin stands at 12%, down from 14% in Q2.
"""
chunks = text_splitter.split_text(report_text)
for i, chunk in enumerate(chunks):
print(f"Chunk {i}: {chunk}")
Strategy 2: Structuring Tables for Embedding
Tables are particularly difficult because their meaning is derived from their grid structure. A simple text dump of a table renders it unreadable. The best approach is to convert tables into a semi-structured format, such as JSON or HTML, preserving row and column headers. This allows the embedding model to understand that "Revenue" is associated with the values in the column below it.
When using tools like PyPDF2 or Tabula-py, extract the table as a DataFrame and then convert it into a list of strings where each string represents a row with its headers:
import pandas as pd
# Assume 'df' is a pandas DataFrame extracted from a PDF table
df = pd.DataFrame({
'Category': ['A', 'B'],
'Q1 Revenue': [100, 200],
'Q2 Revenue': [150, 250]
})
# Convert each row to a context-rich string
table_contexts = []
for index, row in df.iterrows():
# Create a readable string for the embedding model
row_string = " | ".join([f"{col}: {val}" for col, val in row.items()])
table_contexts.append(f"Row {index} Data: {row_string}")
print(table_contexts)
# Output: ['Row 0 Data: Category: A | Q1 Revenue: 100 | Q2 Revenue: 150', ...]
Strategy 3: HTML-Based Preprocessing
If your source documents are HTML, leveraging a robust parser like BeautifulSoup can help maintain layout integrity. By analyzing the DOM tree, you can reorder content to reflect visual reading order rather than DOM order, which is critical for multi-column layouts.
from bs4 import BeautifulSoup
import html5lib
# Parse HTML to respect DOM structure
soup = BeautifulSoup(html_content, "html5lib")
# Extract text while preserving structural hints
for div in soup.find_all('div', class_='column'):
# Process column content sequentially
print(div.get_text(separator=' ', strip=True))
Conclusion
Handling complex layouts is not just a preprocessing step; it is a core component of RAG architecture success. By implementing semantic chunking, structuring tabular data into readable formats, and utilizing robust HTML parsers, developers can significantly improve the quality of their embeddings. These techniques ensure that the LLM retrieves precise, context-aware information, ultimately delivering more accurate and reliable answers to end-users.