Seymur Kafkas

I work on AI products end to end: training and serving models, making them fast and cheap to run, and building the infrastructure and tooling around them.
I'm most at home where research meets systems: computer vision and LLMs on one side, low-level optimization (down to CUDA), backend, and production infrastructure on the other. I studied Computer and Electrical Engineering at Istanbul Technical University.
At work
Most of my work is on the AI platform behind Remini (a gen-ai image/video editing app with 120M+ monthly active users). A few of the bigger things I've built:
- Autoscaler: the system that places models across a large GPU fleet, balancing cost, latency, and demand so everything stays available without overspending.
- Self-hosted AI CI: a self-hosted CI for our ML stack, so builds, tests, and releases run a lot faster.
- Deployment platform: an automated deploy system the team uses to ship models and services, with monitoring and alerts built in.
- CUDA kernels: custom CUDA kernels for our image effects, profiling the hot paths and rewriting them to run faster and cheaper.
- AI features: computer-vision models I've trained and shipped behind user-facing features, like reworking a photo's subject, lighting, and expression.
- Internal tools: an internal app the team uses to launch training jobs, try out open and closed models, generate content, and create, store, and browse image datasets.
I usually build these end to end, from infrastructure and backend to the frontend of the internal tools, so I work fullstack.
Now
I'm at Bending Spoons, where I spent most of my time on Remini, as well as the internal AI platform and tooling. I recently moved to AOL as an AI Researcher.