560 blogs tracked4,950 posts indexed

Rag

We track 15 posts about Rag from 10 engineering blogs. Most active: Pamela Fox, Nic Raboy, Red Hat. Latest post: Oct 6, 2026.

Raw tag behind this topic:rag

Companies writing about Rag

Recent posts

  • How Jev can help improve the efficiency of RAG pipelines (opens on the source site)

    ThoughtWorks ·

    After generation, use a yes-or-no for each claim. Split the LLM's answer into individual claims. A simple sentence splitter is usually enough. For each claim, ask whether the retrieved chunks support it, then flag or remove anything below your threshold before the answer reaches the user. Why Jev? A grounding check only helps if you run it on every answer, and that's only realistic when each check is fast and cheap. In every case, Jev will deliver a probability and an answer. That's what makes the next part possible. Use the probability, not just the answer The real value of Jev here isn't…

  • RAG Complexity Is a Bet Against the Model (opens on the source site)

    Timescale ·

    One Postgres table with two indexes beat every published RAG pipeline on MuSiQue. Every architectural addition is a bet the model stays weak.

  • RAG Citation Verification: Building Deterministic Byte-Span Validators in TypeScript (opens on the source site)

    SitePoint ·

    Eliminate hallucinated LLM references by building high-throughput, deterministic byte-span citation validators and grounding assertion middleware in TypeScript. Continue reading RAG Citation Verification: Building Deterministic Byte-Span Validators in TypeScript on SitePoint.

  • Should you read the code, is RAG dead, and did Skills kill MCP? (opens on the source site)

    GitHub Old ·

    We dive into these questions and other AI hot takes on the latest episode of the GitHub Podcast. The post Should you read the code, is RAG dead, and did Skills kill MCP? appeared first on The GitHub Blog.

  • Orchestrate production RAG with OpenShift AI (opens on the source site)

    Red Hat ·

    In the previous post in this series, we built a streaming retrieval-augmented generation (RAG) pipeline that parses, chunks, embeds, and writes to Milvus in a single Ray Data script. It works well. But it is a monolithic script. When parsing fails at file 847 of 1,000, you rerun everything from scratch. The post Orchestrate production RAG with OpenShift AI appeared first on Red Hat Developer.

  • RAG in Go: A Vulnerability Research Tool (opens on the source site)

    William Kennedy ·

    Introduction In the previous post, you saw how you can use tools to add information to an LLM query. In this post, we’ll see another method of adding information to an LLM called RAG, or Retrieval-Augmented Generation. The idea of RAG is that you want the LLM to have access to information that wasn’t available to it when it was initially trained. You do it by storing documents in your own database along with their embedding. I won’t go into the technical details of embedding, but think of it as a way to convert a piece of text into a vector. The magic is that if two pieces of text have…

  • Graph RAG: Elevating AI with Dynamic Knowledge Graphs (opens on the source site)

    Stack Abuse ·

    Introduction In the rapidly evolving landscape of Artificial Intelligence, Retrieval-Augmented Generation (RAG) has emerged as a pivotal technique for enhancing the factual accuracy and relevance of Large Language Models (LLMs). By enabling LLMs to retrieve information from external knowledge bases before generating responses, RAG mitigates common issues such as hallucination

  • RAG Demystified: From Math to Self-Hosted Code (opens on the source site)

    RisingStack ·

    In today’s AI hype you cannot miss the term “RAG,” which stands for Retrieval Augmented Generation. In plain English, it stands for customizing large language model reasoning with your own context and knowledge. I searched a lot of resources and AI-generated content for this fairly simple technique to be explained well. I’m still looking for […] The post RAG Demystified: From Math to Self-Hosted Code appeared first on RisingStack Engineering.

  • GPT-5: Will it RAG? (opens on the source site)

    Pamela Fox ·

    table.evalresults { background: #FAFAFA; /* pale neutral for table body */ width: 100%; border-collapse: collapse; } table.evalresults th, table.evalresults td { border: 1px solid #ddd; padding: 8px; } table.evalresults th { background-color: #DDEAF2; /* accessible blue header */ text-align: left; } table.evalresults caption { background-color: #F7E7DA; /* optional: caption strip */ padding: 6px; font-weight: bold; } blockquote.evalprompt { background: #F5F9FC; /* very pale blue for good contrast */ border-left: 4px solid #4A90E2; /* accessible blue accent */ margin: 1em 0; padding: 0.75em…

  • GPT-5: Will it RAG? (opens on the source site)

    Pamela Fox ·

    OpenAI released the GPT-5 model family today, with an emphasis on accurate tool calling and reduced hallucinations. For those of us working on RAG (Retrieval-Augmented Generation), it's particularly exciting to see a model specifically trained to reduce hallucination. There are five variants in the family: gpt-5 gpt-5-mini gpt-5-nano gpt-5-chat: Not a reasoning model, optimized for chat applications gpt-5-pro: Only available in ChatGPT, not via the API As soon as GPT-5 models were available in Azure AI Foundry, I deployed them and evaluated them inside our popular open source RAG template. I…

  • Red-teaming a RAG app: gpt-4o-mini v. llama3.1 v. hermes3 (opens on the source site)

    Pamela Fox ·

    When we develop user-facing applications that are powered by LLMs, we're taking on a big risk that the LLM may produce output that is unsafe in some way - like responses that encourage violence, hate speech, or self-harm. How can we be confident that a troll won't get our app to say something horrid? We could throw a few questions at it while manually testing, like "how do I make a bomb?", but that's only scratching the surface. Malicious users have gone to far greater lengths to manipulate LLMs into responding in ways that we definitely don't want happening in domain-specific user…

  • Red-teaming a RAG app: gpt-4o-mini v. llama3.1 v. hermes3 (opens on the source site)

    Pamela Fox ·

    When we develop user-facing applications that are powered by LLMs, we're taking on a big risk that the LLM may produce output that is unsafe in some way - like responses that encourage violence, hate speech, or self-harm. How can we be confident that a troll won't get our app to say something horrid? We could throw a few questions at it while manually testing, like "how do I make a bomb?", but that's only scratching the surface. Malicious users have gone to far greater lengths to manipulate LLMs into responding in ways that we definitely don't want happening in domain-specific user…

  • How to Make a RAG Application With LangChain4j (opens on the source site)

    Nic Raboy ·

    Retrieval-augmented generation, or RAG, introduces some serious capabilities to your large language models (LLMs). These applications can answer questions about your specific corpus of knowledge, whil... The post How to Make a RAG Application With LangChain4j appeared first on DEV.

  • Evaluating gpt-4o-mini vs. gpt-3.5-turbo for RAG applications (opens on the source site)

    Pamela Fox ·

    The azure-search-openai-demo repository was first created in March 2023 and is now the most popular RAG sample solution for Azure. Since the world of generative AI changes so rapidly, we've made many upgrades to its underlying packages and technologies over the past two years. But we've never changed the default GPT model used for the RAG flow: gpt-35-turbo. Why, when there are new models that are cheaper and reportedly better, such as gpt-4o-mini? Well, changing the model is one of the most significant changes you can make to impact RAG answer quality, and I did not want to make the change…

  • Evaluating gpt-4o-mini vs. gpt-3.5-turbo for RAG applications (opens on the source site)

    Pamela Fox ·

    The azure-search-openai-demo repository was first created in March 2023 and is now the most popular RAG sample solution for Azure. Since the world of generative AI changes so rapidly, we've made many upgrades to its underlying packages and technologies over the past two years. But we've never changed the default GPT model used for the RAG flow: gpt-35-turbo. Why, when there are new models that are cheaper and reportedly better, such as gpt-4o-mini? Well, changing the model is one of the most significant changes you can make to impact RAG answer quality, and I did not want to make the change…

Related topics

This page is generated automatically from the engineering blogs we follow. Every post links to its source, where it was published. See all sources.

Privacy choices

Reading never requires analytics. These choices last 90 days on this browser.

Essential sign-in and security storage always stays on. Read the privacy notice.