560 blogs tracked4,950 posts indexed

#ai

50 posts · 16 companies · newest first

3

Part 5: Operating an LLM system: observability, cost, routing, and the platform underneath (opens on the source site)

Your service can be 100% up and still quietly approving the wrong things, burning its budget, or failing over into untested quality. Level 5 is the infrastructure that lets you see your decisions, bound your spend, route and fail over between models, kill bad behavior in seconds — and the platform that makes all of it possible.

aibuilding-softwareexcerpt only · body stays at the source
From the web
5

Part 4: Safety and governance for LLM systems: guardrails, PII, audit, and memory (opens on the source site)

The level where an LLM system stops being a demo and earns the right to touch real data and real decisions: layered guardrails that fail closed, PII handled at the boundary, an immutable audit trail, and scoped memory. By the time an LLM system is making decisions that matter, “it usually works” is no longer the bar. This is Level 4 of the maturity model — safety and governance — and it’s where four disciplines that teams tend to bolt on late have to be designed in instead. They share one idea: don’t trust a single point to do the right thing. Layer independent guardrails so a miss at one is…

aibuilding-softwareexcerpt only · body stays at the source
From the web
6

Part 3: Knowing when your agent doesn’t know: the confidence layer (opens on the source site)

The most important number an agent produces isn’t its answer — it’s how sure it is. Compose that number from independent signals, check it’s calibrated, grade the high-stakes calls with a second model, and route the rest to humans well. This is Level 3 of a six-level maturity model for running LLM systems in production. Levels 1 and 2 got you to where the system works and you can see it working. Level 3 is confidence: the system acts on its own only when its calibrated confidence is high, grades the decisions that matter with an independent judge, and routes everything it’s unsure about to a…

aicontributedexcerpt only · body stays at the source
From the web
7

Part 2: Evals as a deployment gate — and how to know when they drift (opens on the source site)

If you can deploy a prompt change without an eval failing the build, you don’t have evals — you have a notebook. And once the gate is green, the slow leaks are still coming for you. Here’s the gate, the baseline, and the shadow-eval loop that catch both. This is Level 2 of the maturity model: evaluation. The principle is short — you don’t ship on hope, you ship on a gate, and then you watch for drift afterward. A gate protects the moment of deploy. Drift detection protects the weeks in between. You need both, and they’re built from different machinery. Most teams “do evals” the way they once…

aibuilding-softwareexcerpt only · body stays at the source
From the web
8

Part 1: Make your AI agents boring: the determinism layer (opens on the source site)

The trick to putting LLM agents in high-stakes systems isn’t a smarter model — it’s containing the model to one node so the rest of the system is ordinary, testable code. Here are the structural moves, with the contracts and types to implement them. Demos love autonomous agents that loop, call tools, and “figure it out.” Production hates them. The moment an agent’s behavior depends on which path the model wandered down today, you can’t test it, can’t audit it, and can’t let it touch anything that matters. In a regulated or high-consequence system — money movement, healthcare, infrastructure —…

aibuilding-softwareexcerpt only · body stays at the source
From the web
9

Part 6: An operating system for coding agents: the disciplined build (opens on the source site)

Coding agents are great for an afternoon and mediocre for a quarter. Here’s the small set of files and rules — with the actual configs — that keeps quality from decaying across months and thousands of edits, whether you run one agent or a fleet of them on the same repo.

aicontributedexcerpt only · body stays at the source
From the web
10

NVIDIA, Microsoft Kick Off a New Beginning for Windows PCs With RTX Spark and AI Agents (opens on the source site)

At a Microsoft event in San Francisco on Wednesday, Jensen Huang and Satya Nadella outlined how NVIDIA and Microsoft are co-engineering hardware and software for AI agents to run on Windows PCs. NVIDIA was founded because of Windows, Huang said. Now AI agents are coming to Windows. “If you look at the entire journey of […]

aiagentic-aiexcerpt only · body stays at the source
From the web
12

A Practical Way to Check Your AI Search Visibility (opens on the source site)

Last month, I built an AI agent that could search the web before answering a question. While working on the article around that project, I got another project idea and that was around the visibility of a website in the world of AI-powered search. The idea essentially revolves around the fact that whether the AI-powered search mentions your website in its answer or not or whether it cites your website as a source or not, and lastly, compare these resulst with the regular Google search results. Does your website show up when AI-powered search answers questions related to the topics it covers?…

aiexcerpt only · body stays at the source
From the web
13

Why Telecom Operators Are Building Their AI Strategy on Open Models (opens on the source site)

Telecom operators are increasingly building their AI strategies on open models — and the reasons go beyond mere cost. Open models give telcos the ability to trust, control and customize AI across their most critical workloads — from autonomous networks to customer care. NVIDIA’s latest State of AI in Telecommunications report reflects this shift, with […]

aisoftwareexcerpt only · body stays at the source
From the web
14

Smarter personalization: How property data helps us understand user price sensitivity (opens on the source site)

Introduction Existing systems estimate user price sensitivity primarily from spending behavior or demographic proxies. They do not systematically account for residential property values, which can indicate a user’s financial circumstances. This omission creates three limitations: Limited property insights: Existing profiles do not account for property values. One-dimensional profiles: Users with similar spending patterns but different living standards receive the same classification. Regional variation: Manual classification does not adapt well to differences between property markets. This…

engineeringanalyticsexcerpt only · body stays at the source
From the web
16

ReviewBench: An open benchmark for AI code review (opens on the source site)

We’re launching ReviewBench, a benchmark for code review agents built on representative GitHub pull requests, multi-source ground truth, calibrated evaluation, and production-aligned metrics. The post ReviewBench: An open benchmark for AI code review appeared first on The GitHub Blog.

ai---mlgithub-copilotexcerpt only · body stays at the source
From the web
17

From Scan to Treatment Plan, AI Helps Close Breast Cancer’s Deadliest Gaps (opens on the source site)

Breast cancer is the most commonly diagnosed cancer among American women — yet the gaps in care are wide. A majority of women over age 40 skip the recommended annual screening. Radiologists are reading more mammograms with fewer colleagues. And when a diagnosis arrives, the tests that inform treatment can take weeks to return results. […]

aiai-for-goodexcerpt only · body stays at the source
From the web
19

A new, bespoke static site generator to replace Jekyll (opens on the source site)

My blog began as a blosxom (Perl) site running on a VPS. In 2011 I moved to the new GitHub Pages, with the site generated statically by Jekyll (Ruby) on a GitHub server. It was a no-brainer: easier, faster, and cheaper, better in every way. After 15 years of Jekyll, this week I replaced it with a new, custom-built static site generator, dubbed ssg, in “C with templates” C++20. The ~8KLoC source closely follows my personal coding style including templated arenas and slices, zero dependencies, and a libc-free core. It’s wicked fast, and a complete, cold generation of my blog takes 150ms on my…

cppmetaexcerpt only · body stays at the source
From the web
21

NVIDIA DGX Spark 64GB Gives Developers More Ways to Build and Scale Local AI (opens on the source site)

Local AI is becoming more useful by the token. As AI agents move from experiments into everyday development, increasingly capable open models are shrinking to fit on more devices, giving builders more to run locally. Coming this month, NVIDIA DGX Spark will be available with 64GB of unified memory from top manufacturer partners — Acer, […]

aiagentic-aiexcerpt only · body stays at the source 62 comments on HN (opens the discussion)
From the web
23

How NVIDIA GPUs Help Accelerate OpenAI’s GPT-6 Astra Ultrafast (opens on the source site)

GPT-6 Astra Ultrafast, running on NVIDIA Blackwell GPUs, is available now in the OpenAI API and to eligible ChatGPT Work and Codex users. Accelerated by inference optimizations through OpenAI’s models that tap into the capabilities of the NVIDIA Blackwell architecture, Ultrafast offers up to 8x faster token generation than the Astra Standard mode. For developers, […]

aiai-infrastructureexcerpt only · body stays at the source
From the web
26

Is sandboxing sufficient to contain rogue agents? (opens on the source site)

Quick caveats: this is a post on AI safety, written by a cryptography professor. If that troubles you, you should read something else. I try hard not to work on AI (except when the topic occasionally tosses itself in my path), so in this post I’m mostly trying to referee arguments made by others. If … Continue reading Is sandboxing sufficient to contain rogue agents? →

aisecurity-researchexcerpt only · body stays at the source 54100 comments on HN (opens the discussion)
From the web
27

NVIDIA Opens Applications for 2027–2028 Graduate Fellowships With Awards Up to $60,000 (opens on the source site)

Bringing together the world’s brightest minds and the latest accelerated computing technology leads to powerful breakthroughs that help tackle some of the biggest research problems. To foster such innovation, the NVIDIA Graduate Fellowship Program provides grants, mentors and technical support to doctoral students doing outstanding research relevant to NVIDIA technologies. The program, in its 26th […]

aicorporateexcerpt only · body stays at the source
From the web
37

SF Ruby Conference: regular ticket prices end September 30 (opens on the source site)

Authors: Camila Mirabal, Tech writer, and Travis Turner, Tech Editor Topics: Developer Community, Rails, AI Regular SF Ruby Conference tickets are $450 through September 30. Meet Ruby friends, play The Pier, join small-group conversations, and explore the 2026 program in San Francisco. Not a single person in the Ruby community has told us they go to conferences for anything other than meeting new people and catching up with old friends. Our community needs that shared warmth and joy more than ever. But going to a conference alone can be intimidating. You walk into a room full of people you…

developer-communityrailsexcerpt only · body stays at the source
From the web
38

Personalization without user identity (opens on the source site)

How Airbnb uses proximity signals to personalize without relying on individual user history.By: Wei Jiang, Bin Xu, Bharathi Thangamani, Weiwei Guo, Sundar Srinivasavaradhan, Tracy Yu, Huiji Gao, Swapnil Ghike, Michael KinotiGreat personalization starts with knowing your user. But what happens when the user is a stranger?A significant share of Airbnb users arrive without a login, without a recent search history, or without any prior booking — especially those landing from paid advertising or organic search. For these users, the ML models that power search ranking, destination recommendations,…

technologymachine-learningexcerpt only · body stays at the source
From the web
39

Free the models: Harness design at the frontier (opens on the source site)

Model routers are everywhere right now: a small model reads each turn and picks which model should handle it. The inconvenient truth is that the router is always less capable than the model it’s choosing for. Replit Agent lets the model decide instead. The main agent, or core loop, chooses its subagents’ tier and effort, and adjusts its own as the task unfolds. Given that freedom, Astra hands routine implementation to less costly subagents and decides for itself where its tokens are worth spending. On both DeepSWE and Terminal-Bench, Replit Agent is Pareto-efficient against Astra on its own:…

engineeringaiexcerpt only · body stays at the source
From the web
40

We tested our own WAF with frontier AI models. Here’s what we found (opens on the source site)

We built a WAF tester that adapted each request based on what the WAF blocked or passed. This helped us explore variations that a fixed test might miss. We ran it across six attack categories on an authorized staging environment and discovered detection gaps worth fixing. Here’s how the loop worked, what got through, and what we did about it.

aiai-wafexcerpt only · body stays at the source
From the web
43

The State of Kotlin in 2026 Report (opens on the source site)

In 2026, Kotlin turned fifteen. Over those fifteen years, Kotlin has built a mature ecosystem, been adopted by companies around the world, and established a large and active developer community. Fifteen years is long enough to understand which technology trends lasted and which didn’t. This report looks back at Kotlin’s first fifteen years to see […]

newsaiexcerpt only · body stays at the source
From the web
46

Is Your “Human-in-the-Loop” Actually Slowing You Down? Here’s What We Learned (opens on the source site)

In the rush to adopt AI and automation, many teams implement human-in-the-loop (HITL) frameworks. They believe that involving a person in the process solves the problems with reliability, quality, and trust. But as we’ve learned from real engineering workflows and integrations, the story isn’t that easy. In some contexts, humans-in-the-loop do improve outcomes, but in others, they can unintentionally become bottlenecks that limit speed, scalability, and innovation. In this post, we’ll analyze when human-in-the-loop is truly valuable, when it slows systems down, and how to strike the right…

cc-by-sacontributedexcerpt only · body stays at the source
From the web
47

Next.js applications, powered by Vite: introducing Vinext 1.0 (opens on the source site)

Vinext 1.0 graduates from an AI experiment to a production-ready framework, letting developers run Next.js apps on Vite. This release brings advanced cache warming, broader compatibility, and an automated testing pipeline.

aicloudflare-workersexcerpt only · body stays at the source 41 comment on HN (opens the discussion)
From the web
50

Cloudflare’s 2026 Annual Founders’ Letter (opens on the source site)

The Internet is changing more today than at any point since Cloudflare launched back on September 27, 2010. As automated traffic surpasses human activity, we reflect on the rise of AI agents, new creators, and how we can help build a fair, sustainable future for the web.

From the web
50 shown

Privacy choices

Reading never requires analytics. These choices last 90 days on this browser.

Essential sign-in and security storage always stays on. Read the privacy notice.