For all the good and bad that has come with the LLM domination of software development recently, one aspect I can definitely appreciate is that everything feels on the table when it comes to figuring out how to best approach our jobs. Granted, that’s also the part that terrifies me. It’s definitely my and many […] The post Even with LLMs, Let’s Pair Together appeared first on Atomic Spin.
Not every message you send to Claude Code needs the most capable model. A quick question about a Git command runs on the same model as a refactor across three services, unless you remember to switch models first. Jev, a new model from TypeSafe AI, can make that decision for you in a few hundred milliseconds and for a fraction of a cent. So I built jev-router, an open source router that asks Jev which Claude model each of your messages needs, and sends it there. What Jev is After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI?I’ve spent the last 2…
Introduction Merchant reviews contain useful details about food quality, portion size, packaging, and value. Finding these details often requires reading many comments. Aggregate ratings simplify comparisons, but they do not explain what shaped each score. We built the Automated Merchant Review Summary System to turn written reviews into concise summaries. Consumers can scan an overall summary or focus on a specific topic. Merchants can identify what customers value and which parts of the experience need attention. This article focuses on how we generate and maintain consumer-facing review…
In the current AI discourse, particularly when AI is unreliable or misused, it’s often pointed out that AI is simply a tool. This is a fair way to think about it. Like any tool, there are things it’s good for and things it’s not. Some people are better at using it than others. Some folks […] The post In the Age of AI, You Need Structure appeared first on Atomic Spin.
Why we think LLMs can be useful and why we will not replace all of our models with themContextAt Medium, we have many Machine Learning models that we use to label stories automatically. These affect what stories we recommend to readers.Here’s some examples:a few of our text classification models. All diagrams and charts made by the authorSome Clarifications on our Machine Learning policyBefore we go deep on this project, I just wanted to clarify a few things about how we stand regarding AI in general.Medium has been training internal models with user and post data for a long time now. We…
Expedia Group Technology — EngineeringCombine a GraphQL schema, hints and a LLM to generate contextual GraphQL mock responsesPhoto by Samuel Vazquez somewhere in New ZealandA product developer on our navigation header team spent an afternoon hand-typing a 200-line GraphQL JSON mock response so they could keep building the UI while a resolver was pending to be implemented. The schema changed the next morning. We threw the mock away. That is the boring tax on every GraphQL prototype: mocks that drift, fixtures that rot, frontends blocked on backends, and demos slipping because nobody wanted to…
Authors: Ying Li, Arjun Rao, Shradha SehgalIntroductionRecommendations sit at the heart of the Netflix experience. Our current production models rely on thousands of hand‑crafted features over users, items, and interactions, along with specialized architectures for sequence modeling, feature interactions, and multi‑task objectives. This stack has evolved over many years to support diverse content types (movies, series, games, live, podcasts) and product surfaces, but its complexity makes it costly to onboard new use cases: adding a content type or surface can require significant feature…
netflixgenaiexcerpt only · body stays at the source
Introduction The efficacy of semantic search relies on the accuracy of the underlying Knowledge Graph (KG). In high-velocity domains like on-demand food delivery or e-commerce, the catalog of entities like dishes, products, and merchants changes rapidly. Current methods for KG construction and maintenance face three critical challenges: Inaccuracy and hallucination from Large Language Models (LLMs): Automated models often infer relationships based on statistical text co-occurrence rather than semantic reality. For instance, an LLM might incorrectly classify “Pho” as a child of “Italian Noodle…
How Airbnb teams build trustworthy Generative AI products by treating evaluation as a first-class engineering discipline; not an afterthought.Nestled into the lush hillside, this stunning modern retreat features striking natural wood architecture, terraced balconies, and a serene landscape.By: Rohit Girme, Dan Miller, Mia Zhao, Lifan Yang, Clint KellyIntroductionGenerative AI breaks a lot of the assumptions that used to hold true for software testing. Unlike traditional software, LLM outputs are non-deterministic, and “correct” is subjective. Because so much judgment is involved, you often…
Part 1: From one support bot to a framework At Grab, AI agents have evolved from interesting team prototypes into production services used every day by millions of merchants, drivers, and consumers. Today, more than 500 services run on our internal agent framework, over 50 Model Context Protocol (MCP) servers are registered on our remote MCP framework, and a single Large Language Model (LLM) gateway fronts every model call across the company, handling billions of tokens each month. None of this was designed up front. It began as the plumbing behind one internal support bot, which then…
By AI Platform’s Model Runtime team and Inference teamIntroductionMost organizations consume LLMs through hosted APIs. Netflix went further — we run the full stack ourselves, from model deployment through inference, inside our existing production environment rather than a separate ML silo. Some of those decisions weren’t obvious, and a few revealed their trade-offs only under production load.This post focuses on the choices where alternatives were seriously considered: engine selection, model packaging, API surface design, deployment strategy, and output constraints enforcement. The goal is…
Training an LLM is the easy part. The hard part is designing experiments and evaluations that you can trust enough to know whether the new model is actually an improvement.By: Baharak SaberidokhtIntroductionShipping a production LLM system means iterating fast on improvements to something that is, by construction, non-deterministic. Models drift, judges disagree with themselves, references regenerate as different strings, and bugs may persist until the next release, because retraining takes weeks. Most of this friction comes from infrastructure challenges, not model quality, and the fixes…
technologyaiexcerpt only · body stays at the source
Expedia Group Technology — InnovationUsing large language models to reveal bottlenecks in Spark SQL execution plansPhoto by Luis del RíoIf you’ve ever stared at a 300-plus-node physical plan at 2 a.m. trying to spot a missing broadcast or one cursed skewed partition, this is for you.Spark makes it deceptively easy to write complex SQL that looks correct but quietly turns into a performance and cost problem at scale. A query that runs fine on day one can slow to a crawl as data grows, joins get wider, and aggregations become more nested. Suddenly, jobs take hours instead of minutes, clusters…
Türkçe yapay zekâ teknolojilerinin gelişimine katkı sağlayacak önemli bir Ar-Ge adımını paylaşmaktan mutluluk duyuyoruz.VNGRS ve BtcTurk | Teknoloji olarak, Sanayi ve Teknoloji Bakanlığı programı kapsamında başvurusunu gerçekleştirdiğimiz “Kumru VLM — Görsel Dil Modeli ve Belge İşleme” projemiz desteklenmeye hak kazandı.13 Haziran 2026 tarihinde gerçekleştirilen imza töreninde, VNGRS Genel Müdürü Aydın Han ve BtcTurk | Teknoloji CEO’su Deniz Oktar’ın katılımıyla proje iş birliğimizi resmileştirdik. Bu kapsamda; Türkçe için sıfırdan geliştirilecek, metin ve görsel verileri birlikte anlayabilen…
Before working for 2 years on the Apache APISIX API gateway, I was mainly oblivious to API gateways. It’s only by working with them that I understood their value. Decoupling the client and the server unlocks a lot of options: moving authentication to the API Gateway, securing APIs, deduplicating API requests, etc. In this post, I want to describe how the same pattern applies to AI. AI gateways AI gateways work in a similar way.
Authors: Ahmed Ben Yahmed, Antoine Schnepf, Karim Kassab, and Mélissa Tamine.The 14th International Conference on Learning Representations (ICLR 2026) was held from April 23 to 27, 2026, at the Riocentro Convention and Event Center in Rio de Janeiro, Brazil. It was the first time the conference made its way to South America. As one of the three flagship venues in machine learning, ICLR gathers each year several thousand researchers and practitioners from academia and industry around the latest advances in representation learning.This edition was particularly vibrant, with a strong emphasis on…
Author: Fabian HöringAgentic systems powered by LLMs can be incredibly impressive in demos. With a few well-crafted prompts, they can demonstrate reasoning, calling tools, and solving complex tasks [1]. Demos are effective at showcasing what’s possible. Production environments, however, are where those capabilities are tested at scale and under real-world conditions.The same agent that performs perfectly on curated examples can behave unpredictably when exposed to real users. Inputs vary widely, conversations evolve over multiple turns, and small prompt changes can lead to unexpected…
We used DSPy to turn prompt engineering for our relevance judge into a measurable, automated optimization loop, improving task performance, cost, and how reliably it works in production.
Key Contributors: Moein Hasani, Hamidreza Shahidi, Trace Levinson, Guanghua ShuIntroductionAt Instacart, we are laser-focused on improving the user experience by making shopping feel easy, engaging, and personalized. Our discovery surfaces play a central role in bringing this to life. Alongside explicit Search intents, discovery is our opportunity to meet customers’ implicit needs, presenting them with the most relevant and inspiring content we have to offer. The main discovery surface within the Instacart app, referred to here as the “Shopping Hub”, is one of the most critical in this…
This artical is a quick introduction into setting up the Gemini API client SDK on Android to perform Optical Character Recognition (OCR) on images. Let’s jump straight in: The Gemini API gives you access to the latest generative AI models from Google. Gemini is a series of multimodal generative AI models developed by Google. Gemini […] The post OCR with Gemini LLM on Android first appeared on Blundell.
Optional Google Analytics helps us understand visits. Microsoft Clarity records masked interactions to improve the site. Optional tools stay off unless you choose them. Privacy details.