Research Engineer

Antonio Lobo Santos

LLM Post-training · Interpretability

I work at DFKI in Saarbrücken on the SOOFI project, as part of the post-training team. My focus is interpretability of new LLM architectures.

Post-training Interpretability LLM Evaluation Model Serving

Saarbrücken, Germany

Antonio Lobo-Santos

Research & Engineering Focus

Post-training & Interpretability

Post-training open LLMs at DFKI, and studying what new architectures learn and how their internals change during post-training.

LLM Evaluation

Evaluation pipelines for classifiers and LLMs: synthetic data, custom metrics and moral-capability tests, run on HPC clusters.

LLM Serving & HRI

Low-latency vLLM serving with quantised open-weight models, connected to ROS2 robots for human-robot interaction studies.

Production Software

Two years building Python/TypeScript microservices, Kubernetes deployments and Kafka/NiFi pipelines for monitoring platforms.

About

I'm a research engineer at DFKI in Saarbrücken, where I work on the SOOFI project as part of the post-training team. My focus is interpretability of new LLM architectures.

Before that I was an academic visitor at Oxford, working on learning stability and equilibria in multi-agent systems. At CSIC-IIIA in Barcelona I evaluated the moral capabilities of classifiers and LLMs, and built the LLM serving stack for a socially assistive robot.

I started out as a software engineer at Axión in Seville, working full-time while studying Mathematics and Computer Engineering.

Research & Engineering Experience

Research Engineer - LLM Post-training & Interpretability

DFKI, German Research Center for Artificial Intelligence

Sept 2026 – Present Saarbrücken, Germany

Member of the post-training team in the SOOFI project, which builds open foundation models. I work on interpretability of new LLM architectures.

Academic Visitor - Multi-Agent Systems & Game Theory

University of Oxford

March 2026 – July 2026 Oxford, UK

Worked on game-theoretic conditions for learning stability and equilibria in multi-agent systems.

Research Intern → AI Research Engineer - LLMs, Robotics & Evaluation

CSIC-IIIA, Spanish National Research Council

Nov 2024 – Feb 2026 Barcelona, Spain
  • Built moral-capability evaluations for text classifiers and LLMs, based on Moral Foundations Theory and Schwartz's theory of values.
  • Generated synthetic datasets and designed metrics for robustness and value-alignment tests.
  • Fine-tuned BERT-style classifiers with custom losses; ran experiments on HTCondor and multi-GPU servers.
  • Built the vLLM serving pipeline (quantised open-weight models) for real-time HRI studies.
  • Integrated the dialogue components with the ROS2 robot stack, including ASR/TTS.

Lecturer in Probabilistic Graphical Models

University Pompeu Fabra

Jan 2025 – April 2025 Barcelona, Spain

Taught probabilistic graphical models, with a focus on Bayesian inference.

Undergraduate Researcher - LLMs & Mathematical Knowledge Graphs

University of Seville

Sept 2023 – July 2024 Seville, Spain

Built a LangChain/RAG pipeline that extracts mathematical knowledge from LaTeX into OWL knowledge graphs. This became my bachelor's thesis and a journal paper.

Software Engineer - Platform, Microservices & Real-Time Monitoring

Axión

July 2021 – July 2023 Seville, Spain

Full-time while completing dual B.Sc. degrees in Computer Engineering and Mathematics.

  • Built Python/FastAPI and TypeScript/NestJS microservices for monitoring platforms used in transport, ports and public spaces.
  • Set up an on-prem high-availability MicroK8s cluster and migrated the company's web apps to it with GitLab CI/CD.
  • Ran Kafka and NiFi ETL pipelines; built real-time features with WebSockets, GraphQL and Elasticsearch.
  • In the second year, led a small team and mentored two interns.

Selected Projects

MariChatmen - written Andalûh language model

Qwen3.5 adapted to written Andalusian Spanish: tokeniser expansion, continued pretraining, SFT and ORPO.

LLM fine-tuningDialectal AIHugging Face

Moral-capability evaluation for LLMs and classifiers

Evaluation of moral reasoning in BERT-style classifiers and LLMs, using synthetic data and custom metrics.

Manuscript in preparation

Real-time LLM serving for HRI

vLLM serving of quantised open-weight models for a ROS2 robot, with structured dialogue and ASR/TTS.

vLLMROS2HRI

Deep Open-Set Recognition framework

PyTorch Lightning/Hydra framework for open-set recognition with ResNet, EfficientNet, ViT and DINOv2 backbones.

View Repository

PPO/GRPO continuous-control RL framework

PPO and critic-free GRPO implemented from scratch for continuous control.

View Repository

Flood-IDSS

Flood decision-support app combining data sources, ML models, an API and dashboards.

View Repository

Mathematical Knowledge Graphs with LLMs

LLM/RAG pipeline that turns LaTeX mathematics into OWL knowledge graphs.

Read Publication

AB Data Challenge - Finalist

Transformer-based anomaly detection on utility time-series data.

Time seriesTransformersFinalist

ShadowPals

Shadow-puppetry learning app using MediaPipe hand landmarks; I wrote the hand-pose parametrisation.

View Repository

Earlier Projects & Awards

App Inventor Against Cyberbullying

Mobile application recognised as MIT App Inventor "App of the Month" for community cyberbullying awareness.

Read Article

Smarthuerto

Award-winning smart agriculture app that won the EC2CE Grow-Lab contest and was featured by MIT App Inventor.

Read Article

Publications & Talks

Publications

ACM/IEEE HRI 2026 Demo

EMY: Supporting Autism Therapy with a Socially Assistive Robot

Sara Cooper, Antonio Lobo-Santos, et al.

HRI '26: 21st ACM/IEEE International Conference on Human-Robot Interaction

I worked on the LLM dialogue and system integration.

View Article
Modelling, 2025

Enhancing Mathematical Knowledge Graphs with Large Language Models

A. Lobo-Santos, J. Borrego-Díaz

Modelling (MDPI), based on my bachelor's thesis

View Article
ECAI 2025 Workshop

Trustworthy AI Through Dual-Role Reasoning

Workshop contribution

Dual-role reasoning as a way to evaluate and improve the safety of LLM reasoning.

Current Research

Interpretability of new LLM architectures

At DFKI, within the SOOFI post-training team.

Moral-capability evaluation for classifiers and LLMs

Manuscript in preparation.

NLP pragmatics and formal grammars

Pragmatics problems studied with context-free grammar formalisms.

Invited Talks

October 2025 · Madrid, Spain

Invited Speaker: LLM-Tools for Coding

La Moncloa (Spanish Government)

How to use LLM coding tools well, and where they go wrong.

November 2025 · Barcelona, Spain

Seminar: LLM-Tools for Coding

CSIC-IIIA

Seminar for IIIA researchers on prompting and working with coding assistants.

View Seminar Details

Earlier Publications

Sept 2023

Building the "mapamático" (map-matic)

Antonio Lobo-Santos, Pablo Martín Berná, Juan Núñez Valdés

Epsilon Journal of Mathematics Education, Number 114

View PDF
Mar 2023

Infinity: Some Curiosities and its Teaching in High School

Antonio Lobo-Santos, Pablo Martín Berná, Juan Núñez Valdés

Números Journal, Volume 113

View Article

Education & Training

M.Sc. in Artificial Intelligence

UPC–UB–URV

2024 – June 2026 Barcelona, Spain

Honours in Machine Learning, Deep Learning, Reinforcement Learning and Complex Networks.

B.Sc. in Mathematics

University of Seville

2019 – 2024 Seville, Spain

Probability, statistics, optimisation and analysis.

B.Sc. in Computer Engineering

University of Seville

2019 – 2024 Seville, Spain

Graduated with 9 honour distinctions, including Algorithms and Complexity, Computer Architecture, Computer Networks, and Intelligent Systems.

Bachelor's Thesis: Large Language Models and Knowledge Graphs to Model Mathematical Knowledge

Read Thesis

Certifications

BlueDot Logo

AI Alignment Course

BlueDot Impact • November 2024

View Credential

Technical Skills

ML / AI

PyTorchTransformersPost-trainingInterpretabilityBERT-style classifiersRepresentation learningRAGLangChainLangGraphDSPyModel evaluationSynthetic dataCustom metrics

LLM serving / inference

vLLMQuantisationOpen-weight LLMsBatchingGPU server deploymentStructured outputsPydantic schemas

RL / optimisation

PPOGRPOSACTD3Reward modellingContinuous-control experimentsMulti-agent learning dynamics

Robotics / HRI

ROS2Socially assistive robotsReal-time interaction pipelinesASR/TTS integrationStructured dialogue systems

Infrastructure / systems

DockerKubernetesMicroK8sGitLab CI/CDLinuxSlurmHTCondorMulti-GPU serversObservability

Backend / data

PythonFastAPITypeScriptNestJSGraphQLWebSocketsPostgreSQLArangoDBKafkaApache NiFiElasticsearch/KibanaRabbitMQ

Languages

Spanish Native
English Advanced (C1)
French Intermediate (B2)
Catalan Basic