Split, Croatia — available for new work

Generative models that reach production.

Luna AI Lab is an AI research studio that also helps companies with software and process automation. We design and train generative and probabilistic models, then take them the rest of the way: distributed training, evaluation you can defend, and a system that actually runs.

  • arXiv:2601.03753
  • 2 US AI patents
  • Kaggle Expert, top 1%
  • 10 Q1 papers, ~300 citations
01

The practice

Three things I do well enough to charge for. Everything else I will tell you to hire elsewhere.
Automatizacija poslovnih procesa za tvrtke u Hrvatskoj →

Generative modelling

Diffusion, flow matching and functional generative networks for high-dimensional conditional sampling. When iterative denoising is too slow or too expensive to run, I collapse it into single-shot generation without losing ensemble calibration, fidelity or diversity.

  • Diffusion
  • Flow matching
  • FGN
  • One-shot generation
  • Ensemble calibration
  • CRPS / REV

LLM and agent systems

Fine-tuning at 70B scale with PEFT/LoRA and 8-bit quantisation, retrieval that returns the right context instead of merely nearby text, and multi-agent setups that hold together under self-play. Typed outputs, and an evaluation harness before anything reaches a user.

  • PEFT / LoRA
  • bitsandbytes
  • FAISS
  • Chroma
  • ColBERT
  • LangChain
  • DSPy

ML in production

A model that only runs in a notebook is not finished. Distributed multi-GPU training, containerised inference, serverless GPU deployment, monitoring and retraining. I have taken a paid product through this entire stack on my own, billing included.

  • PyTorch DDP
  • Docker
  • RunPod
  • GCP Cloud Run
  • xarray / dask
  • WandB
  • Hydra
02

Selected work

Problem, what got built, and the number it moved.

  1. GEM-1, GEM-2, GEM-3

    Salient Predictions · 2023–2026

    Co-led the technical design and training of the diffusion and flow-matching models behind GEM-1, then GEM-2 (equal contribution): a ~275M-parameter one-shot extension of diffusion models built on functional generative networks. It jointly models global atmospheric dynamics and the variables people actually decide on — daily maximum and minimum temperature, precipitation, wind extremes — in a single forward pass. I owned the pipeline end to end: architecture, training data, distributed multi-GPU training, evaluation systems and deployment. Co-author on GEM-3, now in production.

    Outperforms NOAA GEFS and ECMWF IFS ENS on CRPS skill at 1–40 day leads, with state-of-the-art relative economic value and positive skill maintained to 126 days. Two US AI patents.

  2. transcribevoice.app

    Solo product · live

    A paid audio-transcription service, built alone. WhisperX plus NeMo speaker diarisation packaged as a multi-stage Docker image on RunPod serverless GPUs; Next.js on Vercel; Firebase for auth and Firestore; Cloud Storage for audio and transcripts; Stripe Checkout for credits; Cloud Functions running the yt-dlp and ffmpeg transcode pipeline.

    Research, infrastructure, frontend, payments and support — one person, one stack, in production.

  3. Menstrual health prediction

    Bellabeat (Y Combinator) · 2022–2023

    A multi-task encoder-decoder transformer over cycle history, symptoms, mood, heart rate, skin temperature and respiratory rate, unified into one pipeline. Shipped to users end to end, from data engineering and GCP monitoring through to how the predictions were presented in the app.

    Period start MAE 2.3 days, period end MAE 0.68 days — roughly 50% better than the previous baseline — and ovulation F1 0.922.

  4. Kaggle competitions

    Expert · rank 551 of 205,000

    Unfamiliar domain, fixed deadline, top-percentile result — repeatedly, and across modalities that share almost nothing. Lumbar spine MRI classification 15th of 1,874. EEG seizure detection 54th of 2,767. Accelerometer sleep-state time series 42nd of 1,877. Vesuvius 3D segmentation 39th of 1,249.

    The useful signal is not any single placing. It is that medical imaging, EEG, time series and 3D vision all came out in the same band.

  5. Numerai

    Classical ML, deployed

    CatBoost and XGBoost regression on encrypted, era-structured financial features with multi-target training, era-aware hyperparameter sweeps and out-of-fold evaluation. Daily inference runs unattended on Cloud Run behind an HTTP trigger.

    Proof that not every problem needs a neural network — and that the deployment story matters more than the model choice.

03

How this works

Small, fixed first step. No retainer before there is evidence.

  1. Diagnostic sprint

    Fixed scope, fixed price, one to two weeks. You bring the problem and whatever historical data exists. You get back a backtested model, an honest evaluation, and a written answer to whether this is worth building at all. Sometimes that answer is no, and this is the cheap way to find out.

  2. Build

    Architecture, training, evaluation and deployment of the real system. We agree the metric before I start, and you see progress against it every week rather than at the end.

  3. Keep it alive

    Once it runs: retraining, drift monitoring, and the changes that follow from what production teaches you. Monthly, and cancellable.

Not in scope 24/7 on-call, on-premise or SCADA integration work, and staff augmentation. If that is what the project actually needs, I will say so on the first call rather than three months in.

04

Contact

Describe the problem in a few sentences. I reply to everything that is not a template.

viktor.cikojevic@luna-ai-lab.com