LLM
We track 95 posts about LLM from 51 engineering blogs. Most active: Red Hat, Etsy, SitePoint. Latest post: Oct 9, 2026.
Companies writing about LLM
Recent posts
From Activity to Intent: Generating User Journeys with LLMs (opens on the source site)
Lin Zhu | Sr. Staff Machine Learning Engineer; Manan Kalra | Machine Learning Engineer II; Logan Jeon | Sr. Machine Learning Engineer; Ye Liu | Staff Machine Learning Engineer; Xiangyi Chen | Sr. Machine Learning Engineer; Jaewon Yang | Principal Machine Learning Engineer; Jinwen Xu | Manager II, Machine Learning Engineering; Tingting Zhu | Sr. Manager, Engineering; Sudarshan Lamkhede | Director, Machine Learning EngineeringPinterest is built to get inspired and then turn the inspiration into realization — a dinner, a renovation, a wedding, a new skill. That only works if we understand more…
Production-grade LLMs and agents: a field guide (opens on the source site)
A maturity model for taking agents from an impressive demo to a system people can depend on — with a self-assessment and the map to a deep-dive on each layer.
Part 5: Operating an LLM system: observability, cost, routing, and the platform underneath (opens on the source site)
Your service can be 100% up and still quietly approving the wrong things, burning its budget, or failing over into untested quality. Level 5 is the infrastructure that lets you see your decisions, bound your spend, route and fail over between models, kill bad behavior in seconds — and the platform that makes all of it possible.
Guide to the OWASP LLM Top 10 (opens on the source site)
Large language models (LLMs) have become mainstream. Recent estimates suggest there are over 1.1 billion ChatGPT users. LLMs and AI agents have become a part of every aspect of our lives, from writing emails to coding and everything in between. Yet, this scenario is a cybersecurity incident waiting to happen. Many developers aren’t used to ...
Part 4: Safety and governance for LLM systems: guardrails, PII, audit, and memory (opens on the source site)
The level where an LLM system stops being a demo and earns the right to touch real data and real decisions: layered guardrails that fail closed, PII handled at the boundary, an immutable audit trail, and scoped memory. By the time an LLM system is making decisions that matter, “it usually works” is no longer the bar. This is Level 4 of the maturity model — safety and governance — and it’s where four disciplines that teams tend to bolt on late have to be designed in instead. They share one idea: don’t trust a single point to do the right thing. Layer independent guardrails so a miss at one is…
Backdoors in LLMs: Why model scanning isn't enough (opens on the source site)
Red Hat ·
You pull a model from Hugging Face. Maybe you merge in a LoRA. You run benchmarks, spot-check a few completions, and ship it. That workflow assumes the model you tested is the one you'll get in production. The post Backdoors in LLMs: Why model scanning isn't enough appeared first on Red Hat Developer.
Mitigating memorization in LLMs (opens on the source site)
The following is part of a series of posts about 2026 summer intern projects – for more, see “What the interns have wrought, special jumbo 2026 edition”
Quiz: LLM Application Development With Python (opens on the source site)
In this quiz, you’ll revisit the core concepts covered in the LLM Application Development With Python learning path. You’ll check your understanding of connecting to LLM APIs, crafting effective prompts, working with LangChain, adding retrieval-augmented generation (RAG) with embeddings and vector databases, building AI agents with Pydantic AI and LangGraph, and connecting LLMs to external tools with the Model Context Protocol (MCP). [ Improve Your Python With 🐍 Python Tricks 💌 – Get a short & sweet Python Trick delivered to your inbox every couple of days. >> Click here to learn more and…
Building LLM Primitives from Scratch in TypeScript: Tokenizers, Vectors, and Attention (opens on the source site)
You will understand and implement the foundational mechanics of LLMs—including BPE tokenization, vector search, and self-attention math—entirely in zero-dependency TypeScript. Continue reading Building LLM Primitives from Scratch in TypeScript: Tokenizers, Vectors, and Attention on SitePoint.
Even with LLMs, Let’s Pair Together (opens on the source site)
For all the good and bad that has come with the LLM domination of software development recently, one aspect I can definitely appreciate is that everything feels on the table when it comes to figuring out how to best approach our jobs. Granted, that’s also the part that terrifies me. It’s definitely my and many […] The post Even with LLMs, Let’s Pair Together appeared first on Atomic Spin.
Route every Claude Code message to the right model with Jev (opens on the source site)
Pulumi ·
Not every message you send to Claude Code needs the most capable model. A quick question about a Git command runs on the same model as a refactor across three services, unless you remember to switch models first. Jev, a new model from TypeSafe AI, can make that decision for you in a few hundred milliseconds and for a fraction of a cent. So I built jev-router, an open source router that asks Jev which Claude model each of your messages needs, and sends it there. What Jev is After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI?I’ve spent the last 2…
Automated merchant review summary system with integrated feedback (opens on the source site)
Grab ·
Introduction Merchant reviews contain useful details about food quality, portion size, packaging, and value. Finding these details often requires reading many comments. Aggregate ratings simplify comparisons, but they do not explain what shaped each score. We built the Automated Merchant Review Summary System to turn written reviews into concise summaries. Consumers can scan an overall summary or focus on a specific topic. Merchants can identify what customers value and which parts of the experience need attention. This article focuses on how we generate and maintain consumer-facing review…
Trust, but verify: Atomic claim checking against LLM hallucinations (opens on the source site)
Elastic ·
Hallucinations kept an LLM out of our knowledge base cleanup for years. Splitting generation from atomic claim verification fixed it and cleared a four-year backlog of duplicate articles.
Build Deterministic LLM Eval Suites in CI with Vitest and Zod (opens on the source site)
Build deterministic, zero-cost CI evaluation suites for LLM applications in TypeScript using Vitest, replay fixtures, and Zod schema contracts. Continue reading Build Deterministic LLM Eval Suites in CI with Vitest and Zod on SitePoint.
Build a Cost-Aware LLM Router in Node.js with Claude Opus 5.5 and GPT-6 Sol (opens on the source site)
Build a TypeScript LLM router in Node.js that balances Claude Opus 5.5 and GPT-6 Sol with token budgeting, latency circuit breakers, and tier fallbacks. Continue reading Build a Cost-Aware LLM Router in Node.js with Claude Opus 5.5 and GPT-6 Sol on SitePoint.
I don't like LLMs (opens on the source site)
I have a lot of mixed feelings about AI and LLM technology. I’m fascinated by its effect on our profession, excited by the potential gains in productivity - and thus the products we could rapidly build. On the other hand, I’m fearful of the damage AI might cause: agent swarms taking over our virtual and physical infrastructure, designing bio weapons. But, back on my first hand, LLMs might also design miracle cures, and come up with clever ways to raise our prosperity. Fundamentally I don’t think we have a choice about riding on the AI technology train. It’s a wild ride and I just hope we’ll…
Catching poor LLM performance and accuracy before deployment (opens on the source site)
Red Hat ·
vLLM is the leading open source inference and serving engine for LLMs. Averaging 64 commits a day with bi-weekly releases, the open source project goes through a significant amount of code changes rapidly. Red Hat AI Inference offers an integrated inference platform powered by vLLM, llm-d and vLLM's llm-compressor. The post Catching poor LLM performance and accuracy before deployment appeared first on Red Hat Developer.
Modeling LLM Context Costs and Tiered Pricing Beyond 200k Tokens (opens on the source site)
Calculate nonlinear LLM context pricing tiers, prompt caching break-evens, and context compaction trade-offs across major foundation model APIs. Continue reading Modeling LLM Context Costs and Tiered Pricing Beyond 200k Tokens on SitePoint.
Build Hybrid Automated Code Review with Deterministic Rules and LLMs (opens on the source site)
Construct a two-tier GitHub Actions code review workflow that combines fast deterministic linters and AST rules with structured semantic LLM evaluators. Continue reading Build Hybrid Automated Code Review with Deterministic Rules and LLMs on SitePoint.
Understanding W8A8 INT8 LLM quantization: Accuracy and performance results (opens on the source site)
Red Hat ·
In Understanding W8A8 INT8 LLM quantization: Half the size, better performance, same accuracy, we compressed a Llama 3.1 8B Instruct model from 14.9 GB to 8.0 GB using 8-bit integer (INT8) W8A8 quantization with SmoothQuant and Generative Pre-trained Transformer Quantization (GPTQ). The post Understanding W8A8 INT8 LLM quantization: Accuracy and performance results appeared first on Red Hat Developer.
How to Orchestrate Multi-Call Conversations with an LLM and Twilio Conversation Memory in Python (opens on the source site)
Twilio ·
Learn how to build a Python FastAPI voice agent that uses Twilio Conversation Memory to remember callers across separate phone calls, so if someone hangs up and calls back, the agent picks up right where the conversation left off.
How to Orchestrate Multi-Call Conversations with an LLM and Twilio Conversation in Node.js Memory (opens on the source site)
Twilio ·
Learn how to orchestrate multi-call voice conversations using Node.js, OpenAI, and Twilio Conversation Memory to persist caller context across separate calls.
How to Orchestrate Multi-Call Conversations with an LLM and Twilio Conversation Memory with PHP (opens on the source site)
Twilio ·
In this tutorial you'll make a PHP service using Open Swoole that retains caller context, preferences and action history across multiple separate inbound calls.
SchemaCrawler LLM Context: Extract and Prune Relational DB Schemas (opens on the source site)
Extract, filter, and compact relational database catalogs using SchemaCrawler to feed precise, token-efficient database context to LLM agents. Continue reading SchemaCrawler LLM Context: Extract and Prune Relational DB Schemas on SitePoint.
Understanding W8A8 INT8 LLM quantization: Half the size, better performance, same accuracy (opens on the source site)
Red Hat ·
Large language models are expensive to serve. A model like Llama 3.1 8B in Bfloat16 (BF16) precision occupies roughly 15 GB of GPU memory. In BF16, each of the 8 billion parameters takes 2 bytes to store, which adds up to roughly 15 GB for the weights—and that's not all. The GPU needs memory for the key-value (KV) cache to store context for active requests, alongside intermediate tensor outputs (activations, as we call them) generated during inference. The post Understanding W8A8 INT8 LLM quantization: Half the size, better performance, same accuracy appeared first on Red Hat Developer.
Project HydraFusion: Frontier quality via multi-model orchestration (opens on the source site)
In controlled offline evaluations, HydraFusion’s selective coding workflows matched or exceeded the evaluated Opus 5 baseline while reducing estimated workflow cost. Now available as a research preview in GitHub Copilot. The post Project HydraFusion: Frontier quality via multi-model orchestration appeared first on The GitHub Blog.
What is bring your own LLM (BYO LLM)? (opens on the source site)
Twilio ·
Bring your own LLM (BYO LLM) lets you choose the large language model powering your AI agents rather than accepting whatever your platform ships with.
Speeding up LLM inference with P-EAGLE in vLLM Speculators (opens on the source site)
Red Hat ·
P-EAGLE (Parallel EAGLE), a new speculative decoding algorithm developed by Amazon, brings the next evolution of speculative decoding to Speculators by extending EAGLE-3 with parallel drafting. The post Speeding up LLM inference with P-EAGLE in vLLM Speculators appeared first on Red Hat Developer.
Evaluating LLM guardrail configs locally with EvalHub (opens on the source site)
Red Hat ·
This is part 2 in a series about local guardrail development and evaluation. In my first article, I discussed how to design and develop a guardrail configuration on a local machine, and then tried some manual testing. The post Evaluating LLM guardrail configs locally with EvalHub appeared first on Red Hat Developer.
LLM quantization guide: How to do it, and how it helps (opens on the source site)
Red Hat ·
Model sizes have roughly doubled every year (see figure 1), and GPU memory hasn't come close to keeping up. So when a new frontier open model is released, how do you actually deploy it to serve one user, or perhaps one thousand at a time, on hardware you can realistically get your hands on? The post LLM quantization guide: How to do it, and how it helps appeared first on Red Hat Developer.
Related topics
This page is generated automatically from the engineering blogs we follow. Every post links to its source, where it was published. See all sources.