AI Engineer · Cottbus, Germany

I build AI systems that get used.

Two years on the backend of an industrial-IoT AI platform at Perinet, plus my own deep work on RAG, agents and LLMOps — judged by whether it measurably works.

Available full-time from summer 2026 · open to relocate
01

About

// who & what

Production first. Python/Go services on real-time MQTT streams, FastAPI, Docker/K8s, CI/CD — two years at Perinet, including the model-benchmarking that shaped the platform.

Depth through my own projects. Hybrid RAG with RAGAs + MLflow eval, multi-agent LangGraph systems, an LLMOps tracer, and a GraphRAG stack.

Not only LLMs. A deterministic wind-farm installation scheduler over 20 years of hourly wind — the model only handles the human edges.

M.Sc. AI at BTU Cottbus (thesis phase) · open to AI/ML/LLM Engineer roles across Germany.

M.Sc.
Artificial Intelligence, BTU
~3 yrs
professional software
2 yrs
AI engineering at Perinet
10
projects shipped
stack
PythonLangGraphFastAPIQdrantDockerKubernetesPyTorchMLflowRAGAsGoReactPostgreSQLMQTTGitHub Actions
02

Experience

// where I've worked
Jul 2024 — May 2026

AI Engineer (Working Student)

Perinet GmbH · Cottbus
  • Python & Go services linking LLM workflows to real-time MQTT streams — FastAPI, Docker/K8s, GitHub Actions.
  • Led containerization of the AI platform (chatbot, anomaly detection, sensor explorer).
  • Owned model benchmarking — retrieval speed vs generation quality across variants.
Oct 2021 — May 2022

Software Engineer Trainee

Cognizant · India
  • Enterprise banking apps (COBOL/JCL/DB2) in agile sprints.
03

Publications

// peer-reviewed & preprints
First author Accepted at DSD 2026 · arXiv:2607.02612 [cs.CV] · July 2026

Fusion: A Framework for Unified Sequential Token Adaptation in Vision Transformers

Aravind Pradeep, Samira Nazari, Mahdi Taheri, Christian Herglotz

Token merging, early exiting and token pruning each cut a Vision Transformer's inference cost alone, but combined they destabilise the intermediate representations. Fusion stages them so they cooperate — merge first, check confidence, prune only what continues — with routing modules that adapt compression per input and expose the accuracy/latency trade-off at inference time, no retraining.

48% less inference energy · up to 4× lower calibration error · ImageNet-1k, DeiT-S
04

Projects

// selected open source
Featured

Wind-Farm Installation Planner

Weather-aware scheduler for a 12-turbine build: 48 campaigns × 20 years of hourly wind, priced as risk.

PythonNumPypandasPydanticFastAPIOptimization

Crane lifts can't be paused, have wind limits, and must fit working hours. The scheduling maths stays deterministic; the LLM handles only the human edges.

€5.59M → €3.56M
189 → 64 days
19/19 weather-years

Deutsch-Tutor

AI German tutor that builds your curriculum from your own mistakes.

32 CEFR lessons · error-driven drills · live app
Next.jsTypeScriptSSE streamingZodTailwind

Every error you make in live conversation is logged and turned into drills: the reply streams immediately while a parallel corrector marks exact error spans, and a decay-weighted formula — not an LLM — picks what you practice next.

GraphRAG Agent

Multi-hop answers by traversing a built knowledge graph, with citations.

knowledge graph · k-hop retrieval · cited answers
GraphRAGnetworkxPydanticLLM extractionPython

Entity and relation extraction builds a typed graph first, so a question that spans three documents is answered by walking edges instead of hoping one chunk contains everything.

GraphRAG Studio

Full-stack app: upload docs, watch the graph build, chat over it.

live graph viz · k-hop retrieval · cited chat
Next.jsTypeScriptReactFastAPIGraphRAG

An interactive force-graph front end over the GraphRAG Agent core, so the retrieval path behind every cited answer is something you can actually see.

Multi-Agent Research Pipeline

Planner→Researcher→Writer→Critic on LangGraph, question to sourced report.

LangGraph · strict role boundaries · CI/CD
LangGraphCrewAIPydanticFastAPIDocker

Each role gets a Pydantic-typed contract and no access to the others' tools, which is what keeps a four-agent loop from collapsing into one agent doing everything badly.

RAG Evaluation System

Hybrid BM25+dense+RRF retrieval with a RAGAs/MLflow regression harness.

0.94 hit@5 · 0.96 citation presence
QdrantRAGAsMLflowLangChainFastAPI

Every retrieval change runs against a fixed question set and logs to MLflow, so a quality regression shows up as a failed run rather than a complaint months later.

LLMOps Observability Dashboard

Self-hosted tracing of latency, tokens and cost per model call.

full-stack · multi-stage Docker · 12 tests
FastAPIReactTypeScriptPostgreSQLDocker

Traces land in PostgreSQL rather than a third-party service, which matters when the calls being traced carry customer data.

Multilingual News NLP Pipeline

Whisper → NER → event classification → summary, on a 4 GB GPU.

+13% F1 · 8.4× faster inference
WhisperXLM-RoBERTaPyTorchMLflow

The whole German-language chain was engineered to fit one consumer GPU, which forced quantisation and batching decisions that a bigger card would have hidden.

LLM Fine-Tuning — JD Extractor

QLoRA Qwen2-0.5B extracting structured JSON from job descriptions.

100% JSON validity · <4 min on 4 GB GPU
QLoRAQwen2PEFTPyTorch

Training roughly 0.44% of the parameters was enough to get reliable structured output from messy postings — a small tuned model beating a large prompted one on this task.

Resume Tailor

CLI that tailors résumé + cover letter and self-checks before the PDF.

multi-stage LLM pipeline · self-checking · private repo
PythonLLM agentsLaTeXCLI

ATS and regression checks run before LaTeX compiles, so a fabricated metric or a broken claim fails the build instead of reaching a recruiter.

05

Skills

// toolkit

AI & Agents

LangGraphLangChainCrewAIRAGAgentic AIPrompt EngineeringStructured OutputsMCPTool Use

LLMOps & Evaluation

RAGAsMLflowLangfuseLLM-as-JudgeHybrid RetrievalQdrantChromaDBpgvector

Programming & Backend

PythonGoFastAPIREST / gRPCPyTorchReactTypeScriptSQL / PostgreSQL

Infrastructure & DevOps

DockerKubernetesGitHub ActionsGitLab CI/CDMQTTAzureAWSLinux
06

Education

// academic

M.Sc. Artificial Intelligence

BTU Cottbus
Oct 2022 — 2026 (thesis phase)

Thesis: content-aware Vision Transformer optimisation for edge inference.

English C1German B1Malayalam native

B.Sc. Computer Application

BVM Holy Cross College, Kottayam
2018 — 2021

Foundations in software development, data structures, databases and systems.

07

Contact

// get in touch

Open to AI / ML / LLM Engineer roles across Germany — remote, hybrid or on-site.