I’m a Data Engineer with 6+ years of professional experience, currently working at Entual GmbH in Germany. My professional foundation is in Ab Initio and enterprise-scale data engineering, and my work has expanded into distributed systems, runtime behaviour, GPU computing, local AI, and performance-oriented software.
I’m especially interested in systems where data movement, execution models, memory, and hardware constraints directly shape performance and reliability.
Experimental runtime work focused on translating CUDA-oriented workloads toward Apple Metal, exploring runtime compatibility, execution models, and GPU portability on Apple Silicon.
C/C++ · CUDA · Metal · Apple Silicon · Runtime Systems · GPU Computing
A systems-focused project exploring local execution, tooling, and resource-aware compute workflows with an emphasis on practical infrastructure and performance-conscious design.
Rust · Systems Programming · Local Compute · Performance Engineering
My active development branch on llama.cpp, with work around persistent and file-backed KV cache, mmap-based memory optimisation, Vulkan/runtime improvements, GPU-backed cache behaviour, and server extensions.
C/C++ · LLM Inference · KV Cache · mmap · Vulkan · GPU Computing
Ab Initio and Data Engineering have been the core of my professional engineering career.
I have worked with enterprise-scale data-processing environments where correctness, reliability, scalability, recoverability, operational stability, and performance are critical.
Key areas include:
Ab Initio · GDE · Conduct>It · Control Center · Express>It · TRW
ETL / ELT · Data Integration · Data Quality · Batch Processing · Production Support · Performance Tuning
Broader data-engineering tools and technologies:
SQL · Oracle · PostgreSQL · Python · Shell / KornShell
I completed an M.Sc. in Data Science, which broadened that engineering foundation into machine learning, research workflows, and data-intensive systems.
My current interests sit at the intersection of data systems and lower-level compute:
- distributed and parallel data processing
- runtime and memory behaviour
- systems programming in Rust and C/C++
- GPU computing and hardware-aware optimisation
- local and resource-efficient AI
- LLM inference and cache architecture
- performance engineering across software and hardware boundaries
A recurring question in my work is: where is the real bottleneck, and what layer is actually responsible for it?
Data Engineering
Ab Initio · ETL / ELT · SQL · Oracle · PostgreSQL · Data Quality
Programming & Systems
Python · Rust · C/C++ · Shell · KornShell
Distributed & Parallel Systems
Kafka · Airflow · Ray · AsyncIO · Parallel Processing
Compute & AI Infrastructure
CUDA · Metal · Vulkan · Apple Silicon · LLM Inference · Local AI
Platforms
Linux · macOS
For a more complete view of my engineering work, projects, and background:
- Portfolio: https://perinban.github.io/portfolio/
- GitHub: https://github.com/Perinban
- LinkedIn: https://www.linkedin.com/in/perinban-parameshwaran/



