Systems that hold up when someone asks how they work.
I build data platforms and AI systems that teams can inspect, reproduce and operate: lakehouses, ML pipelines, retrieval systems and the models behind them.
Lucas Rangel Soares de Souza Senior Data & AI Platform Engineer
years building software, data and AI
12+
projects delivered
140+
research and technical articles published
19
companies served
20+
TRY IT
Live demos you can open, test and question.
Each demo covers one area teams hire for, from generative AI to data governance, running on the public Brazilian datasets published on Kaggle. They describe what the records say; they are not legal advice or a supplier recommendation.
A self-hosted language model with no refusal layer, grounded in knowledge bases: it searches procurement notices, contracts and education spending, writes read-only SQL when the answer is a number, cites every record and keeps the thread of a conversation.
1.25 million procurement notices embedded with a self-hosted model and indexed in pgvector: compare keyword search with search by meaning, side by side, next to the aggregate dashboard.
The catalogue of a public data lake: 428 tables and 33,891 columns from raw to analytics layers, with source, lineage, join keys and the privacy rules applied, searchable in the browser.
Browse the map
OPEN SOURCE
The code behind every demo and research project.
Every demo and research project above, grouped by track, with its repository, published data and stack.
Eleven Brazilian public sources, from the school census to procurement, released as raw, trusted and analytics layers with contracts, a privacy gate and per-file hashes.
31 Kaggle datasets, 428 tables and about 4 billion rows of Parquet. Every file is listed in a SHA-256 manifest, and a clean download is verified against it.
A cited research chat over three public bases: text and vector retrieval for notices, read-only SQL for contracts and spending, and gates that run before any model call.
1,250,335 notice embeddings in pgvector. The 2026-10-01 browser session answered 10 of 10 scripted questions.
Numeric questions across all bases can still fall back to text retrieval; a question router is in progress.
Twelve years, from embedded software to AI platforms.
Ten employers across retail, banking, payments, credit and education. Select one to see what I did there.
2014201620182020202220242026
DEC 2024 – NOW · 1 YR 10 MO
Drogasil
Data and AI engineering · Retail
Data and AI engineering for one of Brazil's largest pharmacy chains: pipelines, models and AI agents delivered in one engineering flow, with governance and data quality built in.
What I did
ELT pipelines with dbt and PySpark, and data prepared for AI applications.
Golden ID, deduplication, governance and AI-assisted cataloguing.
LLM, RAG and vector-database solutions, with GitLab CI/CD for data, models and agents.
Results
Pipelines, models and AI agents evolving together in one integrated engineering flow.
Stronger governance, cataloguing and monitoring of data quality.
dbt
PySpark
Airflow
GitLab CI
LLM
RAG
DEC 2024 – NOW · 1 YR 10 MO
Drogasil
Data and AI engineering · Retail
Data and AI engineering for one of Brazil's largest pharmacy chains: pipelines, models and AI agents delivered in one engineering flow, with governance and data quality built in.
What I did
ELT pipelines with dbt and PySpark, and data prepared for AI applications.
Golden ID, deduplication, governance and AI-assisted cataloguing.
LLM, RAG and vector-database solutions, with GitLab CI/CD for data, models and agents.
Results
Pipelines, models and AI agents evolving together in one integrated engineering flow.
Stronger governance, cataloguing and monitoring of data quality.
What memory-bandwidth arithmetic, speculative decoding, a paired benchmark and an audited grader taught me about serving an open model at interactive speed