<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Data, Engineering, and Beyond]]></title><description><![CDATA[Exploring the intersection of engineering, data, and people — building systems and teams that scale.]]></description><link>https://blog.dativo.io</link><image><url>https://substackcdn.com/image/fetch/$s_!SVdI!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7bfbcde6-87b3-4b8c-9b38-3d1b82408e62_800x800.png</url><title>Data, Engineering, and Beyond</title><link>https://blog.dativo.io</link></image><generator>Substack</generator><lastBuildDate>Tue, 04 Aug 2026 20:32:30 GMT</lastBuildDate><atom:link href="https://blog.dativo.io/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Sergey]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[sergeyenin@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[sergeyenin@substack.com]]></itunes:email><itunes:name><![CDATA[Sergey]]></itunes:name></itunes:owner><itunes:author><![CDATA[Sergey]]></itunes:author><googleplay:owner><![CDATA[sergeyenin@substack.com]]></googleplay:owner><googleplay:email><![CDATA[sergeyenin@substack.com]]></googleplay:email><googleplay:author><![CDATA[Sergey]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[What did your coding assistant actually do? ]]></title><description><![CDATA[A coding assistant might become a problem when it calls tools]]></description><link>https://blog.dativo.io/p/what-did-your-coding-assistant-actually</link><guid isPermaLink="false">https://blog.dativo.io/p/what-did-your-coding-assistant-actually</guid><dc:creator><![CDATA[Sergey]]></dc:creator><pubDate>Mon, 27 Jul 2026 17:28:15 GMT</pubDate><enclosure url="https://images.unsplash.com/photo-1764547167756-e75f7944a29f?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwzMnx8Y29waWxvdHxlbnwwfHx8fDE3ODUxNzMxMjl8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://images.unsplash.com/photo-1764547167756-e75f7944a29f?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwzMnx8Y29waWxvdHxlbnwwfHx8fDE3ODUxNzMxMjl8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://images.unsplash.com/photo-1764547167756-e75f7944a29f?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwzMnx8Y29waWxvdHxlbnwwfHx8fDE3ODUxNzMxMjl8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1764547167756-e75f7944a29f?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwzMnx8Y29waWxvdHxlbnwwfHx8fDE3ODUxNzMxMjl8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1764547167756-e75f7944a29f?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwzMnx8Y29waWxvdHxlbnwwfHx8fDE3ODUxNzMxMjl8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1764547167756-e75f7944a29f?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwzMnx8Y29waWxvdHxlbnwwfHx8fDE3ODUxNzMxMjl8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw"><img src="https://images.unsplash.com/photo-1764547167756-e75f7944a29f?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwzMnx8Y29waWxvdHxlbnwwfHx8fDE3ODUxNzMxMjl8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" width="3840" height="2160" data-attrs="{&quot;src&quot;:&quot;https://images.unsplash.com/photo-1764547167756-e75f7944a29f?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwzMnx8Y29waWxvdHxlbnwwfHx8fDE3ODUxNzMxMjl8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:2160,&quot;width&quot;:3840,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Two people flying in a small airplane&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Two people flying in a small airplane" title="Two people flying in a small airplane" srcset="https://images.unsplash.com/photo-1764547167756-e75f7944a29f?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwzMnx8Y29waWxvdHxlbnwwfHx8fDE3ODUxNzMxMjl8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1764547167756-e75f7944a29f?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwzMnx8Y29waWxvdHxlbnwwfHx8fDE3ODUxNzMxMjl8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1764547167756-e75f7944a29f?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwzMnx8Y29waWxvdHxlbnwwfHx8fDE3ODUxNzMxMjl8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1764547167756-e75f7944a29f?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwzMnx8Y29waWxvdHxlbnwwfHx8fDE3ODUxNzMxMjl8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Photo by <a href="https://unsplash.com/@horizonflights">Horizon flights</a> on <a href="https://unsplash.com">Unsplash</a></figcaption></figure></div><p>I gave the real GitHub Copilot CLI - unmodified, straight from GitHub this prompt, which references an MCP (Model Context Protocol) tool server:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;bash&quot;,&quot;nodeId&quot;:&quot;9b382e4a-150f-4619-8774-67314867689e&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-bash">This is a bounded product integration proof. Do exactly this and nothing else:

1. Call release_status from the release-gateway MCP server with this exact run_nonce: &lt;run-nonce&gt;
2. Call release_prepare from the release-gateway MCP server with the same exact run_nonce: &lt;run-nonce&gt;
3. Stop immediately after those two calls and briefly report their outcomes.

Do not run shell commands, read or modify files, call release_publish, retry, delegate,
use subagents, or inspect anything else.
</code></pre></div><p>(`<em><mark data-color="#ffff00" style="background-color: rgb(255, 255, 0); color: rgb(0, 0, 0);">&lt;run-nonce&gt;</mark></em>` is a fresh random value minted for each run &#8212; more on why below.)</p><p>Copilot could see a release-management API exposing release operations over MCP. Between Copilot and everything else stood <a href="https://dativo.io/">Talon</a>, an open-source control plane: one self-hosted Go binary serving an OpenAI-compatible model endpoint and an MCP proxy for tool calls. Copilot kept its native protocol on both.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!I-hr!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac869f2b-ba6b-4cad-8946-1257e5a9a466_1693x929.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!I-hr!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac869f2b-ba6b-4cad-8946-1257e5a9a466_1693x929.png 424w, https://substackcdn.com/image/fetch/$s_!I-hr!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac869f2b-ba6b-4cad-8946-1257e5a9a466_1693x929.png 848w, https://substackcdn.com/image/fetch/$s_!I-hr!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac869f2b-ba6b-4cad-8946-1257e5a9a466_1693x929.png 1272w, https://substackcdn.com/image/fetch/$s_!I-hr!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac869f2b-ba6b-4cad-8946-1257e5a9a466_1693x929.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!I-hr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac869f2b-ba6b-4cad-8946-1257e5a9a466_1693x929.png" width="1456" height="799" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ac869f2b-ba6b-4cad-8946-1257e5a9a466_1693x929.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:799,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1108049,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://blog.dativo.io/i/208334921?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac869f2b-ba6b-4cad-8946-1257e5a9a466_1693x929.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!I-hr!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac869f2b-ba6b-4cad-8946-1257e5a9a466_1693x929.png 424w, https://substackcdn.com/image/fetch/$s_!I-hr!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac869f2b-ba6b-4cad-8946-1257e5a9a466_1693x929.png 848w, https://substackcdn.com/image/fetch/$s_!I-hr!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac869f2b-ba6b-4cad-8946-1257e5a9a466_1693x929.png 1272w, https://substackcdn.com/image/fetch/$s_!I-hr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac869f2b-ba6b-4cad-8946-1257e5a9a466_1693x929.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>What happened: three model turns on gpt-4o-mini, two tool calls, and a stop &#8212; exactly as instructed. The run finished inside its 90-second timeout and left this record (summarized from the signed session):</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;yaml&quot;,&quot;nodeId&quot;:&quot;e88c5f04-67c3-4e9a-b352-cf6ea7663c85&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-yaml">Use case:       coding-assistant
Client:         GitHub Copilot CLI
Provider:       OpenAI
Model:          gpt-4o-mini
Actions:        release_status, release_prepare
Session cost:   $0.003538
Evidence:       5 valid, 0 invalid records</code></pre></div><p>The point of the exercise was never whether Copilot can call tools. It was whether, afterward, the company operating it could answer the questions that arrive with every incident, audit, and invoice:</p><ul><li><p>Which AI use case did the work, and on which model?</p></li><li><p>What data was detected and changed before it reached the provider?</p></li><li><p>Which actions reached the release system &#8212; and did anything reach it that shouldn&#8217;t have?</p></li><li><p>What did the session cost?</p></li><li><p>Can the record be verified after the fact, against a published specification?</p></li></ul><p>One session record answered all five. The rest of this post shows how the experiment worked, what it proves, and &#8212; just as deliberately &#8212; what it does not.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://blog.dativo.io/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://blog.dativo.io/subscribe?"><span>Subscribe now</span></a></p><p></p><h2>What GitHub already gives you &#8212; and where it stops</h2><p>Any honest pitch starts by steel manning the native controls, because they are real. GitHub ships <a href="https://docs.github.com/en/copilot/how-tos/copilot-cli/administer-copilot-cli-for-your-enterprise">enterprise administration for Copilot CLI</a>, MCP server allowlisting policies, <a href="https://github.blog/changelog/2026-02-26-enterprise-ai-controls-agent-control-plane-now-generally-available/">an agent control plane with agent-attributed audit events</a>, and a genuinely useful client-side permission ladder - <em>--allow-tool</em>, <em>--deny-too</em>, trusted directories. My own experiment uses that ladder.</p><p>But look at where each control enforces. The org policies toggle features; the permission flags live in client configuration on each laptop; the audit events tell you an agent acted, in GitHub&#8217;s products , and GitHub&#8217;s own changelog notes that session-activity coverage for the CLI is still on the roadmap. GitHub&#8217;s own documentation is candid about the controls that do <em><strong>not</strong></em> apply to the CLI. What none of these can do, as of this writing, is stand in the request path: inspect the model call that carries your data to the provider, or the tool call that carries an action to your release system. A <a href="https://devopsjournal.io/blog/2026/05/01/Copilot-extension-governance-concerns">practitioner audit of the Copilot extension surfaces </a>found the same pattern from the other side: most surfaces ship with no org-level allowlist, no central gate, and no audit trail.</p><p>Client-side controls govern what a well-behaved client asks for. Nothing in that stack governs what actually crosses the wire.</p><h2>Why the gap is dangerous, not just untidy</h2><p>An agentic coding assistant holds all three ingredients of Simon Willison&#8217;s <a href="https://simonwillison.net/2025/Jun/16/the-lethal-trifecta/">lethal trifecta</a>: access to private data &#128274;, exposure to untrusted content &#9888;&#65039;, and the ability to communicate externally &#128225;. </p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!JMG_!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05607ad4-8a06-4af1-9bd0-0f5d98dd4a20_2092x1046.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!JMG_!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05607ad4-8a06-4af1-9bd0-0f5d98dd4a20_2092x1046.jpeg 424w, https://substackcdn.com/image/fetch/$s_!JMG_!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05607ad4-8a06-4af1-9bd0-0f5d98dd4a20_2092x1046.jpeg 848w, https://substackcdn.com/image/fetch/$s_!JMG_!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05607ad4-8a06-4af1-9bd0-0f5d98dd4a20_2092x1046.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!JMG_!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05607ad4-8a06-4af1-9bd0-0f5d98dd4a20_2092x1046.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!JMG_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05607ad4-8a06-4af1-9bd0-0f5d98dd4a20_2092x1046.jpeg" width="1456" height="728" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/05607ad4-8a06-4af1-9bd0-0f5d98dd4a20_2092x1046.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:728,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;The lethal trifecta (diagram). Three circles: Access to Private Data, Ability to Externally Communicate, Exposure to Untrusted Content.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="The lethal trifecta (diagram). Three circles: Access to Private Data, Ability to Externally Communicate, Exposure to Untrusted Content." title="The lethal trifecta (diagram). Three circles: Access to Private Data, Ability to Externally Communicate, Exposure to Untrusted Content." srcset="https://substackcdn.com/image/fetch/$s_!JMG_!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05607ad4-8a06-4af1-9bd0-0f5d98dd4a20_2092x1046.jpeg 424w, https://substackcdn.com/image/fetch/$s_!JMG_!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05607ad4-8a06-4af1-9bd0-0f5d98dd4a20_2092x1046.jpeg 848w, https://substackcdn.com/image/fetch/$s_!JMG_!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05607ad4-8a06-4af1-9bd0-0f5d98dd4a20_2092x1046.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!JMG_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05607ad4-8a06-4af1-9bd0-0f5d98dd4a20_2092x1046.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Simon Willison&#8217;s lethal trifecta</figcaption></figure></div><p>The tool channel itself is an attack surface - <a href="https://invariantlabs.ai/blog/mcp-security-notification-tool-poisoning-attacks">Invariant Labs demonstrated MCP tool poisoning</a> where the human approves a harmless-looking summary while the model reads a poisoned payload. And asking the model to police itself does not hold: <a href="https://developer.microsoft.com/blog/securing-mcp-a-control-plane-for-agent-tool-execution/">Microsoft&#8217;s own measurement </a> found prompt-only guardrails still leaked a <strong><mark data-color="#ffff00" style="background-color: rgb(255, 255, 0); color: rgb(0, 0, 0);">26.67%</mark></strong> policy-violation rate, which is why their conclusion matches ours &#8212; enforcement has to be deterministic and sit outside the model.</p><p>The operational version of the same gap is quieter but just as real. The model provider has one set of logs, the MCP server another, the client a transcript, finance an account-level invoice, and the data-handling rules live in a policy document. When something goes wrong, there is no single answer to </p><blockquote><p>what did this coding assistant actually do?</p></blockquote><p>and the surface is growing: Anthropic counts <a href="https://www.anthropic.com/news/donating-the-model-context-protocol-and-establishing-of-the-agentic-ai-foundation">more than 10,000 active public MCP servers</a> an assistant can be pointed at.</p><h2>The experiment, step by step</h2><p>Everything below is reproducible from the public <a href="https://github.com/dativo-io/talon-full-demo">talon-full-demo</a> repository; the <a href="https://github.com/dativo-io/talon/blob/main/docs/guides/github-copilot-cli-governance.md">Copilot CLI governance guide</a> walks it end to end.</p><p>Five local components:</p><ol><li><p>GitHub Copilot CLI - the real client, unmodified, run non-interactively.</p></li><li><p>Session shim (:8079) ~55-line reverse proxy that stamps two headers &#8212; session ID and client label</p></li><li><p>Talon gateway (:8080) - authenticates the use case, applies the effective policy, calls OpenAI, signs a record per model call.</p></li><li><p> Talon MCP proxy (:8081)  - filters tool discovery, enforces tool policy, forwards allowed calls, signs a record per tool call.</p></li><li><p>Synthetic release server (:8090)  - stand-in release API with release_status, release_prepare, release_publish; performs no real release.</p></li><li><p>The interesting part was not whether Copilot could use the tools.</p></li></ol><p>Copilot&#8217;s model traffic is redirected with its own standard configuration &#8212; the same custom-endpoint interception pattern now common for governing coding agents. This uses Copilot CLI&#8217;s custom-provider (bring-your-own-key) mode; seat-licensed subscription traffic billed through GitHub&#8217;s own OAuth never transits an interceptable endpoint &#8212; more on that boundary below.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;bash&quot;,&quot;nodeId&quot;:&quot;baefeaa4-f7fb-4eb6-93b1-c01aefdd0233&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-bash">export COPILOT_PROVIDER_TYPE=openai
export COPILOT_PROVIDER_BASE_URL=http://127.0.0.1:8079/v1/proxy/openai/v1
export COPILOT_PROVIDER_API_KEY="$TALON_CODING_ASSISTANT_KEY"
export COPILOT_MODEL=gpt-4o-mini</code></pre></div><p>Note the API key: Copilot receives a Talon agent key representing the <em><mark data-color="#ffff00" style="background-color: rgb(255, 255, 0); color: rgb(0, 0, 0);">coding-assistant</mark></em><mark data-color="#ffff00" style="background-color: rgb(255, 255, 0); color: rgb(0, 0, 0);"> </mark>use case - never the real OpenAI credential. The gateway resolves that key to the use-case identity, computes the effective policy, retrieves the vaulted OpenAI credential itself, and only then calls the provider. A provider key answers <em><strong>which account may call OpenAI</strong></em>; a Talon agent key answers<em><strong> which company AI use case is making this request, and which policy applies</strong></em>.</p><p>Copilot&#8217;s tool traffic gets the mirror treatment &#8212; its MCP config points at Talon&#8217;s proxy, not at the release server:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;json&quot;,&quot;nodeId&quot;:&quot;3ac74afc-276f-4e28-85e5-12ceb47c285b&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-json">{
  "mcpServers": {
    "release-gateway": {
      "type": "http",
      "url": "http://127.0.0.1:8081/mcp/proxy",
      "headers": {
        "Authorization": "Bearer &lt;coding-assistant Talon key&gt;",
        "X-Talon-Session-ID": "copilot-&lt;run-id&gt;",
        "X-Talon-Client": "github-copilot-cli-full-demo"
      }
    }
  }
}</code></pre></div><p>Both planes carry the same session ID &#8212; supplied by the shim on the model plane and the MCP config on the action plane, and recorded as <em><strong>client-asserted</strong></em> attribution: metadata the client supplies under its authenticated agent key, not proof of process identity (more in the boundary section). That is how five records from two different interception points land in one session:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;plaintext&quot;,&quot;nodeId&quot;:&quot;c9643eee-fcda-4edc-abce-d725aa7b56df&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-plaintext">gateway          gpt-4o-mini       allow
proxy_tool_call  release_status    allow
gateway          gpt-4o-mini       allow
proxy_tool_call  release_prepare   allow
gateway          gpt-4o-mini       allow</code></pre></div><p>One more design detail does a lot of work: the synthetic release server appends a receipt for every call that actually reaches it, and every call must carry the run&#8217;s fresh random nonce. Those nonce-stamped receipts are ground truth for upstream effect, independent of anything Copilot &#8212; or Talon &#8212; claims about itself. A stale receipt from yesterday&#8217;s run cannot fake a pass.</p><h2>One record, five answers</h2><p><strong><mark data-color="#ea9999" style="background-color: rgb(234, 153, 153); color: rgb(0, 0, 0);">Which use case, which model?</mark></strong> Every record &#8212; model and tool alike &#8212; carries <em><strong>agent_id: coding-assistant</strong></em>, the session <em><strong>copilot-20260724T132616Z-5a5ba0</strong></em>, and the client label. A request log shows five API calls; this session record shows one company AI use case using three model turns to perform two release actions.</p><p><mark data-color="#f9cb9c" style="background-color: rgb(249, 203, 156); color: rgb(0, 0, 0);">What happened to the data? </mark>The tested run&#8217;s records show email and person entities detected in the input and redacted before the request reached OpenAI. That happened because the request physically transited the control plane - data handling enforced in the path, not described in a policy document. (Honest scope: this covers the input path of intercepted traffic, nothing more.)</p><p><strong><mark data-color="#a4c2f4" style="background-color: rgb(164, 194, 244); color: rgb(0, 0, 0);">Which actions?</mark> </strong>Exactly <em><strong>release_status</strong></em> and <em><strong>release_prepare</strong></em>, each with a signed proxy record and a nonce-matched upstream receipt. No release_publish receipt exists for the run.</p><p><strong><mark data-color="#b4a7d6" style="background-color: rgb(180, 167, 214); color: rgb(0, 0, 0);">What did it cost?</mark></strong></p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;bash&quot;,&quot;nodeId&quot;:&quot;8bf0ae17-a6f6-4902-8409-12f143196b3c&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-bash">Input:       7,031 tokens
Cache read:  32,128 tokens
Output:      123 tokens
Cost:        $0.003538</code></pre></div><p>Not commercially meaningful in itself - the structure is the point. An account-level bill says OpenAI cost $20,000 last month; it cannot say which use cases consumed it, which sessions a spike maps to, or whether those sessions also touched internal systems. And there is a difference in kind, not degree, between reporting and enforcement: a cost dashboard tells you on Tuesday that a session overspent on Monday; a gateway budget denies the *next* model request before provider dispatch, a soft cap that preserves completed work while bounding the overrun at request granularity, demonstrated in a separate n8n scenario in the <a href="https://github.com/dativo-io/talon-full-demo">same demo repository</a>.</p><p><strong><mark data-color="#93c47d" style="background-color: rgb(147, 196, 125); color: rgb(0, 0, 0);">Can it be verified?</mark></strong> That deserves its own section - see it, right below.</p><h3>The evidence is checked</h3><p>The presentation/render is not the proof. The signed records are the proof, and the presentation is a readable projection of them.</p><p>After the run, the session is exported and verified:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;bash&quot;,&quot;nodeId&quot;:&quot;df64971f-ce85-429e-a5fb-f3c36ad0be29&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-bash">talon audit export --format signed-json --session "$SESSION" --output evidence.json
talon audit verify --file evidence.json

Total records: 5
Valid records: 5
Invalid records: 0
Missing signature: 0
Could not parse: 0
Unsupported: 0</code></pre></div><p>The demo&#8217;s presenter re-runs the whole chain before rendering anything human-readable: signatures verified, every record confirmed to belong to this session and this use case, the MCP records asserted to contain exactly the two expected tools, the upstream receipts nonce-checked. It fails closed on any mismatch. Records are signed with HMAC-SHA256 and independently verifiable against <a href="https://github.com/dativo-io/talon/blob/main/docs/reference/evidence-integrity-spec.md">Talon&#8217;s published integrity specification</a>. HMAC is symmetric, so be precise about what that buys: anyone holding the signing key can re-verify every record against the spec alone, which makes the records tamper-evident under the operator&#8217;s key custody - not immutable, and not third-party non-repudiation. Reproduce the run yourself and verify your own records; the pass criteria are public.</p><p>The same signed records feed several audiences from one source &#8212; developers inspect individual records, platform engineering reviews the timeline, FinOps reads spend, security reads data handling and actions &#8212; and they feed Talon&#8217;s compliance report generators (GDPR Article 30 and EU AI Act Annex IV packs) as supporting evidence for reviews, not a compliance determination.</p><h2>Where the boundary ends</h2><p>A control plane must be explicit about what it cannot see. This experiment proves control and evidence for model requests routed through Talon, MCP calls routed through Talon, and tool schemas carried in intercepted model requests. It does <em>not </em>prove control over:</p><ul><li><p>local shell commands and file edits (the demo restricts these with Copilot&#8217;s own <em>--deny-tool</em> flags - client-side safeguards, not Talon enforcement)</p></li><li><p>browser actions </p></li><li><p>direct APIs that bypass Talon</p></li><li><p>what a governed tool did after the call was forwarded - Talon evidences the call that crossed the boundary, not the tool&#8217;s downstream execution</p></li><li><p>the correctness of Copilot&#8217;s generated answer</p></li><li><p>the identity of the operating process independently of its presented credentials</p></li></ul><p>The record describes client provenance as <em><mark data-color="#ffff00" style="background-color: rgb(255, 255, 0); color: rgb(0, 0, 0);">client_asserted:</mark></em> operational attribution, not process attestation. One prerequisite is equally explicit: this pattern requires a client that can point at a custom endpoint with a company-held provider key, which Talon injects from its vault. Seat-licensed subscription traffic, billed through the vendor&#8217;s own OAuth, does not transit the gateway and cannot be governed this way.</p><p>These are not wording details &#8212; they determine which additional controls you still need: endpoint security, repository permissions, sandboxing, workload identity (SPIFFE-style), network restrictions. Treat any vendor&#8217;s non-interception list as an evaluation criterion; a control plane that won&#8217;t show you its boundary is asking you to assume it has none.</p><h2>Rollout: observe, attribute, enforce</h2><p>The way in is not &#8220;govern every AI system in the enterprise.&#8221; It is one existing workflow, three reversible steps:</p><ol><li><p>Observe. &#128065;&#65039;  Point one coding assistant&#8217;s endpoint and MCP config at Talon and change nothing else. Developers keep their client; you get the session timeline you didn&#8217;t have yesterday.</p></li><li><p>Attribute. &#127991;&#65039; Turn on per-use-case identity and session cost. FinOps gets spend per use case and session; security gets data-handling visibility. </p></li><li><p>Enforce. &#128737;&#65039; Scope the tool surface and set session budgets. Actions outside the approved path are denied at the boundary &#8212; with signed denial evidence.</p></li></ol><p>Pilot success criteria, from the run above: developers keep the existing client; model and MCP traffic share one identity and one session; data handling is visible in the record; cost is attributed to the session; every intercepted action is verifiable; no out-of-scope action reaches upstream. First-time setup is about twenty minutes with the <a href="https://github.com/dativo-io/talon/blob/main/docs/guides/github-copilot-cli-governance.md">guide</a>; budget an afternoon if you include evaluation.</p><p>The deliverable that matters is not the first governed workflow - it is that the <em>second</em> one onboards as a config and policy change under an already-reviewed pattern. Security stops approving integrations and starts approving use cases: the review happens once, at the boundary, instead of once per tool.</p><h2>Run it, then tell us what broke</h2><p>You have always been responsible for what your assistant does with your credentials. Now you can hold the record of it.</p><p>&#9889; 60 seconds, no API key: <a href="https://github.com/dativo-io/talon/blob/main/docs/tutorials/quickstart-demo.md">quickstart demo</a> </p><p>&#129517; 20 minutes, the full Copilot scenario: <a href="https://github.com/dativo-io/talon/blob/main/docs/guides/github-copilot-cli-governance.md">governance guide</a>  </p><p>&#129514; The complete experiment: <a href="https://github.com/dativo-io/talon-full-demo">talon-full-demo</a> ) </p><p>&#129413; <a href="https://github.com/dativo-io/talon">Talon</a>  (Apache 2.0)</p><p>And an open invitation: run this scenario against your own workflow - your assistant, your MCP integration, your policy, and co-author the writeup with us, backed by your own signed records. The strongest post in this category won&#8217;t be written by a vendor. It will be written by the first platform team that publishes its evidence.</p>]]></content:encoded></item><item><title><![CDATA[AI transparency isn't a UI problem. It's a distributed systems problem.]]></title><description><![CDATA[The EU AI Act doesn't just introduce new disclosure rules - it exposes a missing architectural layer in enterprise AI.]]></description><link>https://blog.dativo.io/p/ai-transparency-isnt-a-ui-problem</link><guid isPermaLink="false">https://blog.dativo.io/p/ai-transparency-isnt-a-ui-problem</guid><dc:creator><![CDATA[Sergey]]></dc:creator><pubDate>Fri, 17 Jul 2026 12:28:39 GMT</pubDate><enclosure url="https://images.unsplash.com/photo-1701861970884-de7e90db2fa8?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyNHx8bG9uZyUyMGV4cG9zdXJlJTIwb2YlMjBkYXRhJTIwY2VudGVyJTIwc2VydmVyJTIwcmFja3MlMjB3aXRoJTIwYmx1ZSUyMGxpZ2h0JTIwdHJhaWxzJTJDJTIwaW50ZXJjb25uZWN0ZWQlMjBzeXN0ZW1zJTJDJTIwbW9kZXJuJTIwaW5mcmFzdHJ1Y3R1cmUlMkMlMjBjaW5lbWF0aWMlMkMlMjBtaW5pbWFsJTJDJTIwZGFyayUyQyUyMG5vJTIwcGVvcGxlfGVufDB8fHx8MTc4NDI5MDc1NXww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://images.unsplash.com/photo-1701861970884-de7e90db2fa8?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyNHx8bG9uZyUyMGV4cG9zdXJlJTIwb2YlMjBkYXRhJTIwY2VudGVyJTIwc2VydmVyJTIwcmFja3MlMjB3aXRoJTIwYmx1ZSUyMGxpZ2h0JTIwdHJhaWxzJTJDJTIwaW50ZXJjb25uZWN0ZWQlMjBzeXN0ZW1zJTJDJTIwbW9kZXJuJTIwaW5mcmFzdHJ1Y3R1cmUlMkMlMjBjaW5lbWF0aWMlMkMlMjBtaW5pbWFsJTJDJTIwZGFyayUyQyUyMG5vJTIwcGVvcGxlfGVufDB8fHx8MTc4NDI5MDc1NXww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://images.unsplash.com/photo-1701861970884-de7e90db2fa8?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyNHx8bG9uZyUyMGV4cG9zdXJlJTIwb2YlMjBkYXRhJTIwY2VudGVyJTIwc2VydmVyJTIwcmFja3MlMjB3aXRoJTIwYmx1ZSUyMGxpZ2h0JTIwdHJhaWxzJTJDJTIwaW50ZXJjb25uZWN0ZWQlMjBzeXN0ZW1zJTJDJTIwbW9kZXJuJTIwaW5mcmFzdHJ1Y3R1cmUlMkMlMjBjaW5lbWF0aWMlMkMlMjBtaW5pbWFsJTJDJTIwZGFyayUyQyUyMG5vJTIwcGVvcGxlfGVufDB8fHx8MTc4NDI5MDc1NXww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1701861970884-de7e90db2fa8?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyNHx8bG9uZyUyMGV4cG9zdXJlJTIwb2YlMjBkYXRhJTIwY2VudGVyJTIwc2VydmVyJTIwcmFja3MlMjB3aXRoJTIwYmx1ZSUyMGxpZ2h0JTIwdHJhaWxzJTJDJTIwaW50ZXJjb25uZWN0ZWQlMjBzeXN0ZW1zJTJDJTIwbW9kZXJuJTIwaW5mcmFzdHJ1Y3R1cmUlMkMlMjBjaW5lbWF0aWMlMkMlMjBtaW5pbWFsJTJDJTIwZGFyayUyQyUyMG5vJTIwcGVvcGxlfGVufDB8fHx8MTc4NDI5MDc1NXww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1701861970884-de7e90db2fa8?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyNHx8bG9uZyUyMGV4cG9zdXJlJTIwb2YlMjBkYXRhJTIwY2VudGVyJTIwc2VydmVyJTIwcmFja3MlMjB3aXRoJTIwYmx1ZSUyMGxpZ2h0JTIwdHJhaWxzJTJDJTIwaW50ZXJjb25uZWN0ZWQlMjBzeXN0ZW1zJTJDJTIwbW9kZXJuJTIwaW5mcmFzdHJ1Y3R1cmUlMkMlMjBjaW5lbWF0aWMlMkMlMjBtaW5pbWFsJTJDJTIwZGFyayUyQyUyMG5vJTIwcGVvcGxlfGVufDB8fHx8MTc4NDI5MDc1NXww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1701861970884-de7e90db2fa8?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyNHx8bG9uZyUyMGV4cG9zdXJlJTIwb2YlMjBkYXRhJTIwY2VudGVyJTIwc2VydmVyJTIwcmFja3MlMjB3aXRoJTIwYmx1ZSUyMGxpZ2h0JTIwdHJhaWxzJTJDJTIwaW50ZXJjb25uZWN0ZWQlMjBzeXN0ZW1zJTJDJTIwbW9kZXJuJTIwaW5mcmFzdHJ1Y3R1cmUlMkMlMjBjaW5lbWF0aWMlMkMlMjBtaW5pbWFsJTJDJTIwZGFyayUyQyUyMG5vJTIwcGVvcGxlfGVufDB8fHx8MTc4NDI5MDc1NXww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw"><img src="https://images.unsplash.com/photo-1701861970884-de7e90db2fa8?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyNHx8bG9uZyUyMGV4cG9zdXJlJTIwb2YlMjBkYXRhJTIwY2VudGVyJTIwc2VydmVyJTIwcmFja3MlMjB3aXRoJTIwYmx1ZSUyMGxpZ2h0JTIwdHJhaWxzJTJDJTIwaW50ZXJjb25uZWN0ZWQlMjBzeXN0ZW1zJTJDJTIwbW9kZXJuJTIwaW5mcmFzdHJ1Y3R1cmUlMkMlMjBjaW5lbWF0aWMlMkMlMjBtaW5pbWFsJTJDJTIwZGFyayUyQyUyMG5vJTIwcGVvcGxlfGVufDB8fHx8MTc4NDI5MDc1NXww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" width="6000" height="4000" data-attrs="{&quot;src&quot;:&quot;https://images.unsplash.com/photo-1701861970884-de7e90db2fa8?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyNHx8bG9uZyUyMGV4cG9zdXJlJTIwb2YlMjBkYXRhJTIwY2VudGVyJTIwc2VydmVyJTIwcmFja3MlMjB3aXRoJTIwYmx1ZSUyMGxpZ2h0JTIwdHJhaWxzJTJDJTIwaW50ZXJjb25uZWN0ZWQlMjBzeXN0ZW1zJTJDJTIwbW9kZXJuJTIwaW5mcmFzdHJ1Y3R1cmUlMkMlMjBjaW5lbWF0aWMlMkMlMjBtaW5pbWFsJTJDJTIwZGFyayUyQyUyMG5vJTIwcGVvcGxlfGVufDB8fHx8MTc4NDI5MDc1NXww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:4000,&quot;width&quot;:6000,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;a couple of people that are walking down a hallway&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="a couple of people that are walking down a hallway" title="a couple of people that are walking down a hallway" srcset="https://images.unsplash.com/photo-1701861970884-de7e90db2fa8?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyNHx8bG9uZyUyMGV4cG9zdXJlJTIwb2YlMjBkYXRhJTIwY2VudGVyJTIwc2VydmVyJTIwcmFja3MlMjB3aXRoJTIwYmx1ZSUyMGxpZ2h0JTIwdHJhaWxzJTJDJTIwaW50ZXJjb25uZWN0ZWQlMjBzeXN0ZW1zJTJDJTIwbW9kZXJuJTIwaW5mcmFzdHJ1Y3R1cmUlMkMlMjBjaW5lbWF0aWMlMkMlMjBtaW5pbWFsJTJDJTIwZGFyayUyQyUyMG5vJTIwcGVvcGxlfGVufDB8fHx8MTc4NDI5MDc1NXww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1701861970884-de7e90db2fa8?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyNHx8bG9uZyUyMGV4cG9zdXJlJTIwb2YlMjBkYXRhJTIwY2VudGVyJTIwc2VydmVyJTIwcmFja3MlMjB3aXRoJTIwYmx1ZSUyMGxpZ2h0JTIwdHJhaWxzJTJDJTIwaW50ZXJjb25uZWN0ZWQlMjBzeXN0ZW1zJTJDJTIwbW9kZXJuJTIwaW5mcmFzdHJ1Y3R1cmUlMkMlMjBjaW5lbWF0aWMlMkMlMjBtaW5pbWFsJTJDJTIwZGFyayUyQyUyMG5vJTIwcGVvcGxlfGVufDB8fHx8MTc4NDI5MDc1NXww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1701861970884-de7e90db2fa8?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyNHx8bG9uZyUyMGV4cG9zdXJlJTIwb2YlMjBkYXRhJTIwY2VudGVyJTIwc2VydmVyJTIwcmFja3MlMjB3aXRoJTIwYmx1ZSUyMGxpZ2h0JTIwdHJhaWxzJTJDJTIwaW50ZXJjb25uZWN0ZWQlMjBzeXN0ZW1zJTJDJTIwbW9kZXJuJTIwaW5mcmFzdHJ1Y3R1cmUlMkMlMjBjaW5lbWF0aWMlMkMlMjBtaW5pbWFsJTJDJTIwZGFyayUyQyUyMG5vJTIwcGVvcGxlfGVufDB8fHx8MTc4NDI5MDc1NXww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1701861970884-de7e90db2fa8?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyNHx8bG9uZyUyMGV4cG9zdXJlJTIwb2YlMjBkYXRhJTIwY2VudGVyJTIwc2VydmVyJTIwcmFja3MlMjB3aXRoJTIwYmx1ZSUyMGxpZ2h0JTIwdHJhaWxzJTJDJTIwaW50ZXJjb25uZWN0ZWQlMjBzeXN0ZW1zJTJDJTIwbW9kZXJuJTIwaW5mcmFzdHJ1Y3R1cmUlMkMlMjBjaW5lbWF0aWMlMkMlMjBtaW5pbWFsJTJDJTIwZGFyayUyQyUyMG5vJTIwcGVvcGxlfGVufDB8fHx8MTc4NDI5MDc1NXww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Photo by <a href="https://unsplash.com/@mosdesign">mos design</a> on <a href="https://unsplash.com">Unsplash</a></figcaption></figure></div><p>In two weeks, a quiet part of the EU AI Act becomes applicable.</p><p>Most of the discussion around it is about what users should see. Banners. Labels. Watermarks.</p><p>I think that&#8217;s missing the real engineering problem.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://blog.dativo.io/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://blog.dativo.io/subscribe?"><span>Subscribe now</span></a></p><p></p><h2>A simple example</h2><p>Imagine a customer support assistant.</p><p>A user opens your website and starts a conversation. The request passes through an AI gateway, which selects a model based on company policy. The model generates a response. The frontend displays it together with an AI disclosure.</p><p>Everything looks correct.</p><p>Now fast-forward three months. An auditor &#8212; or simply your own security team &#8212; asks a few seemingly straightforward questions:</p><ul><li><p>Which model generated this response?</p></li><li><p>Which policy was active at that moment?</p></li><li><p>Was the AI disclosure actually shown to the user &#8212; and when?</p></li><li><p>Can you prove all of this?</p></li></ul><p>Most organizations can answer parts of these questions.</p><p>Very few can answer all of them with evidence.</p><p></p><h2>We&#8217;ve solved this problem before</h2><p>The instinctive answer is: &#8220;we already log everything.&#8221;</p><p>Logging is necessary. It isn&#8217;t sufficient. Logs describe isolated events &#8212; model selected, request completed, page published, banner rendered &#8212; but they can&#8217;t tell you whether those events belong to the same piece of content.</p><p>We&#8217;ve hit this wall before.</p><p>Microservices turned one request into fifty independent systems.</p><p>The solution wasn&#8217;t more logging.</p><p>It was distributed tracing: one identifier, propagated across every system a request touches.</p><p>AI-generated content is creating the same architectural problem &#8212; except the object being traced isn&#8217;t an HTTP request. It&#8217;s a piece of content, moving through generation, review, publication and translation over weeks instead of milliseconds.</p><p>And this time, correlation isn&#8217;t enough. Traces are sampled &#8212; dropping data is the design. They&#8217;re retained for weeks &#8212; auditors ask about years. Any service can emit spans into any trace &#8212; and rewrite them after the fact &#8212; with no signatures and no identity binding.</p><p>A trace tells you what probably happened, for debugging. This problem needs a record that proves what happened, for accountability.</p><p>That&#8217;s the difference between logging and provenance.</p><h2>The information already exists</h2><p>Every system already knows part of the story. The gateway knows which provider, model and policy were used. The application knows whether disclosure was rendered. The editorial workflow knows who approved the content. The CMS knows when it was published. Individually those records are correct. Collectively they cannot reconstruct the lifecycle.</p><p>The problem isn&#8217;t missing information. It&#8217;s missing continuity.</p><h2>What the law actually asks</h2><p>The quiet part of the AI Act is Article 50, its transparency chapter. It applies from 2 August 2026. Providers must design chatbots so people know they&#8217;re talking to AI. Deployers must label deepfakes and certain AI-generated text. Getting it wrong can cost up to &#8364;15 million or 3% of worldwide turnover &#8212; whichever is higher.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://images.unsplash.com/photo-1697802912100-16c2f7ec814c?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw4MHx8YWklMjBhZ2VudHxlbnwwfHx8fDE3ODQyOTEwNzB8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://images.unsplash.com/photo-1697802912100-16c2f7ec814c?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw4MHx8YWklMjBhZ2VudHxlbnwwfHx8fDE3ODQyOTEwNzB8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1697802912100-16c2f7ec814c?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw4MHx8YWklMjBhZ2VudHxlbnwwfHx8fDE3ODQyOTEwNzB8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1697802912100-16c2f7ec814c?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw4MHx8YWklMjBhZ2VudHxlbnwwfHx8fDE3ODQyOTEwNzB8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1697802912100-16c2f7ec814c?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw4MHx8YWklMjBhZ2VudHxlbnwwfHx8fDE3ODQyOTEwNzB8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw"><img src="https://images.unsplash.com/photo-1697802912100-16c2f7ec814c?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw4MHx8YWklMjBhZ2VudHxlbnwwfHx8fDE3ODQyOTEwNzB8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" width="1080" height="1080" data-attrs="{&quot;src&quot;:&quot;https://images.unsplash.com/photo-1697802912100-16c2f7ec814c?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw4MHx8YWklMjBhZ2VudHxlbnwwfHx8fDE3ODQyOTEwNzB8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1080,&quot;width&quot;:1080,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;a man in a suit with a virtual headset&quot;,&quot;title&quot;:&quot;a man in a suit with a virtual headset&quot;,&quot;type&quot;:&quot;image/jpg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="a man in a suit with a virtual headset" title="a man in a suit with a virtual headset" srcset="https://images.unsplash.com/photo-1697802912100-16c2f7ec814c?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw4MHx8YWklMjBhZ2VudHxlbnwwfHx8fDE3ODQyOTEwNzB8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1697802912100-16c2f7ec814c?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw4MHx8YWklMjBhZ2VudHxlbnwwfHx8fDE3ODQyOTEwNzB8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1697802912100-16c2f7ec814c?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw4MHx8YWklMjBhZ2VudHxlbnwwfHx8fDE3ODQyOTEwNzB8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1697802912100-16c2f7ec814c?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw4MHx8YWklMjBhZ2VudHxlbnwwfHx8fDE3ODQyOTEwNzB8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>But Article 50(2) is the part worth reading twice: generative outputs must be &#8220;marked in a machine-readable format and detectable as artificially generated or manipulated&#8221;, through technical solutions that are &#8220;effective, interoperable, robust and reliable&#8221; as far as technically feasible. Machine-readable. Interoperable. Protocol words, not UI words.</p><p>(One timing nuance: the just-adopted Digital Omnibus gives systems already on the market until 2 December 2026 for the machine-readable marking. The disclosure duties still start on 2 August.)</p><p>One thing Article 50 never says: keep records. The Act&#8217;s logging obligations apply only to high-risk systems - a different chapter entirely. The proof pressure comes from enforcement instead. Market surveillance authorities can demand documentation from any operator. And the EU&#8217;s new Code of Practice on AI-content transparency - finalized in June, voluntary &#8212; makes the trade explicit: sign it, and you get a recognized way to show compliance; don&#8217;t, and you prove to authorities, case by case, that your own measures are good enough.</p><p>The law says: disclose and mark.</p><p>Enforcement asks: can you show that you did?</p><p><strong>The obligations span roles; the evidence spans systems; the systems share no identifiers.</strong></p><h2>The exemption that proves the point</h2><p>A marketing team drafts an article with an LLM. An editor rewrites it. Legal reviews it. The CMS publishes it; a translated version gets syndicated. Months later, somebody asks whether this content originated from AI and whether transparency obligations were met.</p><p>Now read Article 50(4). Its text-disclosure duty is narrow to begin with: it only covers AI-generated text published to inform the public on matters of public interest. And even where it applies, it exempts content that has undergone &#8220;a process of human review or editorial control&#8221; where someone &#8220;holds editorial responsibility for the publication of the content&#8221;.</p><p>Both limbs are questions of fact. What was this published for? Did a human review it &#8212; and who owned that responsibility?</p><p>The legally decisive artifact isn&#8217;t the label. It&#8217;s the evidence.</p><p>The statute never says &#8220;build an evidence chain.&#8221; It just makes relying on its exemptions untenable without one.</p><h2>The missing layer: operational provenance</h2><p>There&#8217;s already a standard for part of this story. C2PA (Content Credentials) gives you asset provenance: a signed history that travels with the file. It matters, and adoption is growing. But manifests are routinely stripped when content crosses platforms, the standard barely covers text - the dominant enterprise output - and by design it says nothing about the operator&#8217;s runtime: which policy was in force, why the router picked that model, whether disclosure duties were met.</p><p>What organizations increasingly need alongside it is operational provenance: a verifiable record of every decision surrounding an AI generation, across every system that touches it.</p><p>Asset provenance travels with the file. Operational provenance survives on the server,  where a C2PA manifest can reference it, not replace it.</p><p>None of the cryptography is new. That&#8217;s a feature. It&#8217;s the same append-only, verifiable-log pattern Certificate Transparency and sigstore proved at internet scale. What doesn&#8217;t exist yet is that pattern applied to AI runtime decisions &#8212; across systems.</p><h2>A control plane for AI operations</h2><p>The next generation of AI infrastructure won&#8217;t stop at routing requests to models. It will treat every generation as an operational event that accumulates evidence over its lifetime:</p><p>AI generation &#8594; runtime policy applied &#8594; disclosure shown &#8594; human review completed &#8594; content published &#8594; evidence exported.</p><p>Almost everything in that chain belongs to the surrounding operational ecosystem, not the model. That&#8217;s why solving transparency exclusively inside an AI gateway will never be enough. Every serious gateway logs today &#8212; that&#8217;s not the gap. The gap is that the gateway&#8217;s record ends at its own boundary. Nothing downstream &#8212; the frontend&#8217;s disclosure, the editor&#8217;s review, the CMS&#8217;s label &#8212; can attach to it.</p><h2>Where Talon fits</h2><p>This is the direction we&#8217;re exploring with <a href="https://dativo.io/">Talon</a>.</p><p>Talon isn&#8217;t trying to become a legal compliance platform. It doesn&#8217;t decide when Article 50 applies. It doesn&#8217;t render disclosures. It doesn&#8217;t watermark images. Those responsibilities belong to applications, CMS platforms and legal teams.</p><p>Talon records the operational provenance of each AI generation: which agent initiated it, which provider and model were selected, which effective policy governed the decision, why routing occurred. Each record is signed and independently verifiable against a published integrity specification.</p><p>The next step is letting downstream systems extend that record with their own evidence. A frontend recording that disclosure was presented. A CMS recording that labeling was applied. An editorial workflow recording that human review completed, the exact fact the text exemption turns on.</p><p>The result isn&#8217;t a prettier dashboard. It&#8217;s a verifiable operational history.</p><p>One honest caveat: attached evidence is attested, not omniscient. The chain proves which system claimed what, when, under which identity  not that every claim is true. That&#8217;s how every audit regime works. What changes is that the claim becomes tamper-evident and durably bound to the generation it describes.</p><h2>Beyond the regulation</h2><p>It&#8217;s tempting to frame all of this around the AI Act. But even without regulation, the same questions keep arriving &#8212; from customers, partners, procurement teams and internal auditors.</p><div class="callout-block" data-callout="true"><p>What happened?</p><p>Who approved it?</p><p>Which policy applied?</p><p>Can you prove it?</p></div><p>Those aren&#8217;t legal questions. They&#8217;re operational questions. Article 50 just makes them harder to ignore.</p><p>A Monday-morning test: pick one AI-generated output that shipped last quarter and reconstruct its full story &#8212; model, policy, disclosure, review, publication &#8212; with evidence, not recollection. If you can&#8217;t, start small. Mint a stable generation ID at the point of generation. Sign the record. Propagate the ID as far downstream as it will go.</p><p>We spent decades building infrastructure to answer one question: what happened to this request?</p><p>AI introduces a different question: what happened to this piece of generated content?</p><p>I suspect the next generation of AI infrastructure will be built around answering it.</p><p>If you&#8217;re solving this differently, I&#8217;d genuinely like to hear how.</p>]]></content:encoded></item><item><title><![CDATA[Fable 5 as the architect, cheaper models as the builders: multi-model coding with predictable costs]]></title><description><![CDATA[The $10 architect and the $2 builders &#8212; a worked session with real rates, the three places predictability leaks, and a hard stop you can prove.]]></description><link>https://blog.dativo.io/p/fable-5-as-the-architect-cheaper</link><guid isPermaLink="false">https://blog.dativo.io/p/fable-5-as-the-architect-cheaper</guid><dc:creator><![CDATA[Sergey]]></dc:creator><pubDate>Mon, 06 Jul 2026 11:48:24 GMT</pubDate><enclosure url="https://images.unsplash.com/photo-1591588582259-e675bd2e6088?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxMDZ8fGFyY2hpdGVjdCUyMGJsdWVwcmludCUyMHdvcmtlcnN8ZW58MHx8fHwxNzgzMzI4NTM4fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Your best model shouldn&#8217;t be typing boilerplate.</p><p>That&#8217;s the whole idea behind the pattern that&#8217;s quietly becoming the default way to run coding agents: let a frontier model - <em>Claude Fable 5</em>, with its 1M-token context read the entire repo, write the plan, and review the result, while cheaper models  <em>Claude Opus 4.8</em>, or <em>Codex</em> on OpenAI&#8217;s side do the fetching, editing, testing, and retrying. Setup guides for Hermes and OpenClaw are already prescribing exactly this split; if your team runs coding agents seriously, some version of it is probably in your dotfiles by autumn.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://images.unsplash.com/photo-1591588582259-e675bd2e6088?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxMDZ8fGFyY2hpdGVjdCUyMGJsdWVwcmludCUyMHdvcmtlcnN8ZW58MHx8fHwxNzgzMzI4NTM4fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://images.unsplash.com/photo-1591588582259-e675bd2e6088?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxMDZ8fGFyY2hpdGVjdCUyMGJsdWVwcmludCUyMHdvcmtlcnN8ZW58MHx8fHwxNzgzMzI4NTM4fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1591588582259-e675bd2e6088?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxMDZ8fGFyY2hpdGVjdCUyMGJsdWVwcmludCUyMHdvcmtlcnN8ZW58MHx8fHwxNzgzMzI4NTM4fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1591588582259-e675bd2e6088?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxMDZ8fGFyY2hpdGVjdCUyMGJsdWVwcmludCUyMHdvcmtlcnN8ZW58MHx8fHwxNzgzMzI4NTM4fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1591588582259-e675bd2e6088?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxMDZ8fGFyY2hpdGVjdCUyMGJsdWVwcmludCUyMHdvcmtlcnN8ZW58MHx8fHwxNzgzMzI4NTM4fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw"><img src="https://images.unsplash.com/photo-1591588582259-e675bd2e6088?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxMDZ8fGFyY2hpdGVjdCUyMGJsdWVwcmludCUyMHdvcmtlcnN8ZW58MHx8fHwxNzgzMzI4NTM4fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" width="4000" height="3000" data-attrs="{&quot;src&quot;:&quot;https://images.unsplash.com/photo-1591588582259-e675bd2e6088?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxMDZ8fGFyY2hpdGVjdCUyMGJsdWVwcmludCUyMHdvcmtlcnN8ZW58MHx8fHwxNzgzMzI4NTM4fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:3000,&quot;width&quot;:4000,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;man in white long sleeve shirt and blue denim jeans standing on white metal ladder&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="man in white long sleeve shirt and blue denim jeans standing on white metal ladder" title="man in white long sleeve shirt and blue denim jeans standing on white metal ladder" srcset="https://images.unsplash.com/photo-1591588582259-e675bd2e6088?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxMDZ8fGFyY2hpdGVjdCUyMGJsdWVwcmludCUyMHdvcmtlcnN8ZW58MHx8fHwxNzgzMzI4NTM4fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1591588582259-e675bd2e6088?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxMDZ8fGFyY2hpdGVjdCUyMGJsdWVwcmludCUyMHdvcmtlcnN8ZW58MHx8fHwxNzgzMzI4NTM4fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1591588582259-e675bd2e6088?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxMDZ8fGFyY2hpdGVjdCUyMGJsdWVwcmludCUyMHdvcmtlcnN8ZW58MHx8fHwxNzgzMzI4NTM4fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1591588582259-e675bd2e6088?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxMDZ8fGFyY2hpdGVjdCUyMGJsdWVwcmludCUyMHdvcmtlcnN8ZW58MHx8fHwxNzgzMzI4NTM4fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Photo by <a href="https://unsplash.com/@markpot123">Mark Potterton</a> on <a href="https://unsplash.com">Unsplash</a></figcaption></figure></div><p>The pitch is cost-effectiveness. And the pitch is right, below is the math with list prices. But there&#8217;s a second promise hiding inside the pattern that almost nobody talks about, and it&#8217;s the one that matters to whoever signs the bill: <strong>done right, this architecture makes agent costs </strong><em><strong>predictable</strong></em>. Decomposable into parts you can reason about, meterable in numbers you can trust, and cappable at the level where the money actually moves the session.</p><p>Here&#8217;s the pattern, the real economics, the three places predictability leaks away, and how to plug them with four HTTP headers and one line of YAML.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://blog.dativo.io/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://blog.dativo.io/subscribe?"><span>Subscribe now</span></a></p><p></p><h2>The pattern in one diagram</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!4rLQ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08e48a78-36ac-4851-b9e1-74c3b055e90f_1440x820.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!4rLQ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08e48a78-36ac-4851-b9e1-74c3b055e90f_1440x820.png 424w, https://substackcdn.com/image/fetch/$s_!4rLQ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08e48a78-36ac-4851-b9e1-74c3b055e90f_1440x820.png 848w, https://substackcdn.com/image/fetch/$s_!4rLQ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08e48a78-36ac-4851-b9e1-74c3b055e90f_1440x820.png 1272w, https://substackcdn.com/image/fetch/$s_!4rLQ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08e48a78-36ac-4851-b9e1-74c3b055e90f_1440x820.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!4rLQ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08e48a78-36ac-4851-b9e1-74c3b055e90f_1440x820.png" width="1440" height="820" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/08e48a78-36ac-4851-b9e1-74c3b055e90f_1440x820.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:820,&quot;width&quot;:1440,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:93265,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://blog.dativo.io/i/205473313?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08e48a78-36ac-4851-b9e1-74c3b055e90f_1440x820.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!4rLQ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08e48a78-36ac-4851-b9e1-74c3b055e90f_1440x820.png 424w, https://substackcdn.com/image/fetch/$s_!4rLQ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08e48a78-36ac-4851-b9e1-74c3b055e90f_1440x820.png 848w, https://substackcdn.com/image/fetch/$s_!4rLQ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08e48a78-36ac-4851-b9e1-74c3b055e90f_1440x820.png 1272w, https://substackcdn.com/image/fetch/$s_!4rLQ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08e48a78-36ac-4851-b9e1-74c3b055e90f_1440x820.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p>The economic logic: </p><p>Planning is few, expensive, high-leverage tokens - you want the smartest model and the biggest context exactly once; </p><p>Execution is many, cheap, repetitive tokens - the transcript gets re-read on every agentic turn, and 90%+ of that re-reading is served from the prompt cache at a tenth of the price.</p><h2>What a session actually costs</h2><p>List prices, per million tokens( actual as for 5th of July, 2026):</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!pV8k!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bbe665c-6a90-4870-9e18-1d1a63859bef_1491x1055.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!pV8k!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bbe665c-6a90-4870-9e18-1d1a63859bef_1491x1055.png 424w, https://substackcdn.com/image/fetch/$s_!pV8k!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bbe665c-6a90-4870-9e18-1d1a63859bef_1491x1055.png 848w, https://substackcdn.com/image/fetch/$s_!pV8k!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bbe665c-6a90-4870-9e18-1d1a63859bef_1491x1055.png 1272w, https://substackcdn.com/image/fetch/$s_!pV8k!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bbe665c-6a90-4870-9e18-1d1a63859bef_1491x1055.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!pV8k!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bbe665c-6a90-4870-9e18-1d1a63859bef_1491x1055.png" width="1456" height="1030" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6bbe665c-6a90-4870-9e18-1d1a63859bef_1491x1055.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1030,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:999581,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://blog.dativo.io/i/205473313?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bbe665c-6a90-4870-9e18-1d1a63859bef_1491x1055.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!pV8k!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bbe665c-6a90-4870-9e18-1d1a63859bef_1491x1055.png 424w, https://substackcdn.com/image/fetch/$s_!pV8k!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bbe665c-6a90-4870-9e18-1d1a63859bef_1491x1055.png 848w, https://substackcdn.com/image/fetch/$s_!pV8k!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bbe665c-6a90-4870-9e18-1d1a63859bef_1491x1055.png 1272w, https://substackcdn.com/image/fetch/$s_!pV8k!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bbe665c-6a90-4870-9e18-1d1a63859bef_1491x1055.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>A realistic very simple feature-build session  - one planning pass over a mid-size repo, two executors working the plan:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!eLXs!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36d086c2-b198-47e4-b481-a49b7a373acf_1491x1055.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!eLXs!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36d086c2-b198-47e4-b481-a49b7a373acf_1491x1055.png 424w, https://substackcdn.com/image/fetch/$s_!eLXs!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36d086c2-b198-47e4-b481-a49b7a373acf_1491x1055.png 848w, https://substackcdn.com/image/fetch/$s_!eLXs!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36d086c2-b198-47e4-b481-a49b7a373acf_1491x1055.png 1272w, https://substackcdn.com/image/fetch/$s_!eLXs!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36d086c2-b198-47e4-b481-a49b7a373acf_1491x1055.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!eLXs!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36d086c2-b198-47e4-b481-a49b7a373acf_1491x1055.png" width="1456" height="1030" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/36d086c2-b198-47e4-b481-a49b7a373acf_1491x1055.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1030,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1042342,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://blog.dativo.io/i/205473313?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36d086c2-b198-47e4-b481-a49b7a373acf_1491x1055.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!eLXs!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36d086c2-b198-47e4-b481-a49b7a373acf_1491x1055.png 424w, https://substackcdn.com/image/fetch/$s_!eLXs!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36d086c2-b198-47e4-b481-a49b7a373acf_1491x1055.png 848w, https://substackcdn.com/image/fetch/$s_!eLXs!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36d086c2-b198-47e4-b481-a49b7a373acf_1491x1055.png 1272w, https://substackcdn.com/image/fetch/$s_!eLXs!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36d086c2-b198-47e4-b481-a49b7a373acf_1491x1055.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p>Three things to mention: </p><p>1. <mark data-color="#cfe2f3" style="background-color: rgb(207, 226, 243); color: rgb(0, 0, 0);">The architect is 65% of the spend on 6% of the calls</mark>. That&#8217;s not waste -  that&#8217;s the design. You paid $4.40 for the one artifact that determines whether the other $2.37 produces working code or thirty loops of confident nonsense. </p><p>2. <mark data-color="#c9daf8" style="background-color: rgb(201, 218, 248); color: rgb(0, 0, 0);">The builders are cheap </mark><strong><mark data-color="#c9daf8" style="background-color: rgb(201, 218, 248); color: rgb(0, 0, 0);">because of the cache</mark></strong>. Over half of all input tokens in this session are cache reads at 0.1&#215; the input rate. The marginal executor turn costs about a tenth of a cent-per-thousand-token intuition suggests. This is what makes execution cost <mark data-color="#c9daf8" style="background-color: rgb(201, 218, 248); color: rgb(0, 0, 0);">linear and boring</mark> &#8212; exactly what you want.</p><p>3. <mark data-color="#d0e0e3" style="background-color: rgb(208, 224, 227); color: rgb(0, 0, 0);">Each part is predictable on its own terms.</mark> </p><p style="text-align: center;">Planner cost &#8776; repo context &#215; rate, a known quantum per task. </p><p style="text-align: center;">Executor cost &#8776; turns &#215; a small, cache-discounted constant. </p><p style="text-align: center;">Decomposable costs are forecastable costs. </p><p>One honest footnote: prompt caches don&#8217;t cross models. Fable&#8217;s cached repo context is not readable by Opus or Codex, each executor builds its own cache on turn one and harvests it for the rest of the session. </p><h3>Does that pattern always win?</h3><p><a href="https://akitaonrails.com/en/2026/07/01/llm-benchmark-sonnet-5-fails-gemini-flash-surprises-sakana-fugu-almost-tier-a/">Fabio Akita&#8217;s widely-cited three-round benchmark</a> concluded that planner&#8211;executor splits are often &#8220;premature optimization&#8221;:  in his tests, solo Opus 4.7 beat every mixed Anthropic combo on quality at comparable cost. It&#8217;s a careful benchmark with a public repo, and if your workload looks like his, solo-frontier is a fine answer. Two things his own data also shows: the one decisive <em>economic</em> win in his rounds was a planner&#8211;executor split (a high-effort GPT planner with a medium executor came in 80&#8211;85% cheaper than solo at near-equal quality), and orchestration paid off exactly where tasks decompose cleanly, which is what large coding tasks increasingly do. But this post doesn&#8217;t need the pattern to win the benchmark. If you run multi-model sessions at all, and the tooling ecosystem is pushing you there, the operational questions below hit you regardless of whether your split beats solo Opus by 5% or loses by 5%.</p><h2>Predictability leaks</h2><p>I built an AI gateway, watched this traffic all day. Predictability dies in three specific places. </p><ul><li><p style="text-align: justify;"><mark data-color="#d9ead3" style="background-color: rgb(217, 234, 211); color: rgb(0, 0, 0);">Leak 1</mark>:<em> your meter doesn&#8217;t understand cached tokens</em>. The two providers report caching in opposite shapes. Anthropic reports cache reads and writes as separate counts - &#8216;input_tokens&#8217; excludes them. At the same time, OpenAI reports &#8216;cached_tokens&#8217; as a subset of &#8216;input_tokens&#8217;. </p></li><li><p style="text-align: justify;"><mark data-color="#d9ead3" style="background-color: rgb(217, 234, 211); color: rgb(0, 0, 0);">Leak 2</mark>: <em>the bill can&#8217;t be attributed</em>. The planner spend lands on your Anthropic invoice; the Codex spend lands on OpenAI&#8217;s. Two dashboards, two currencies of &#8220;usage&#8221;, zero concept that these were one piece of work. </p></li><li><p style="text-align: justify;"><mark data-color="#d9ead3" style="background-color: rgb(217, 234, 211); color: rgb(0, 0, 0);">Leak 3</mark>: <em>nothing stops a runaway session</em>. The industry answer to agent overspend so far is coarse: monthly per-tool caps (<a href="https://finance.yahoo.com/sectors/technology/articles/uber-caps-monthly-employee-ai-180342247.html">Uber&#8217;s much-covered $1,500/tool policy</a>), per-user or per-team budgets (OpenAI&#8217;s enterprise controls, Cloudflare&#8217;s gateway limits). All useful; none match the failure mode. A looping executor doesn&#8217;t blow a month &#8212; it blows a Tuesday afternoon, spread across two providers so that neither one&#8217;s limiter sees the whole picture. To my current knowledge, no popular tool or provider control enforces a budget at the <em>session</em> grain, <em>across providers</em>, with a <em>record of the denial you can hand to finance</em>.</p></li></ul><h2>Plugging the leaks: four headers and one line of YAML</h2><p></p><p>This is the part we built. <a href="https://dativo.io/">Talon</a> is an open-source (Apache-2.0) gateway: point both tools at it, and every request carries the session identity the pattern already implies.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;bash&quot;,&quot;nodeId&quot;:&quot;41d3c4f0-20c6-40e9-8175-3ac78256d7df&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-bash"># planner / executors all send the SAME session id &#8212; that's the entire contract
-H "X-Talon-Session-ID: sess-payments-refactor"
-H "X-Talon-Agent-ID: planner"            # free-form labels, your names
-H "X-Talon-Parent-Agent-ID: orchestrator" # optional parent link</code></pre></div><p></p><p>Claude Code and Codex CLI don&#8217;t even need that &#8212; their native session/subagent headers are recognized automatically. And the cap is one line on the caller:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;yaml&quot;,&quot;nodeId&quot;:&quot;f57a521f-7805-4e58-812c-d831aec545ce&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-yaml">olicy_overrides:
  max_session_cost: 10.00   # soft cap per coding session</code></pre></div><p>What you get is something like this:</p><p><strong>The session becomes one governed unit across both(Anthropic and OpenAI) providers:</strong></p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;markdown&quot;,&quot;nodeId&quot;:&quot;9fb625a9-ac3b-452e-a949-8662ec8ce97e&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-markdown">Session sess-coding-demo
  Requests:  10 (8 allowed, 2 denied, 0 error)
  Providers: anthropic, openai
  Tokens:    in 43 / out 334 / cache-read 14352 / cache-write 840
  Cost:      &#8364;0.012551

  Per-agent:
    generator                 8 req  &#8364;0.011960
    executor  &#8592;generator      2 req  &#8364;0.000591</code></pre></div><p>Cache reads metered as cache reads, priced at cache rates, with the pricing basis recorded in the evidence - leak 1 closed. One session, two providers, per-subagent rollup -  leak 2 closed.</p><p><strong>And when the session hits its cap, it stops, on either provider&#8217;s route:</strong></p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;markdown&quot;,&quot;nodeId&quot;:&quot;a3ee452c-412b-4cdc-88f6-c3de78522acb&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-markdown">request 6 &#8594; HTTP 403
session_budget_exceeded: session spend 0.01 + estimate 0.01 exceeds limit 0.02</code></pre></div><p>The denial isn&#8217;t a log line &#8212; it&#8217;s a signed evidence record carrying the exact numbers the decision was made on <a href="https://github.com/dativo-io/talon/blob/main/docs/guides/governing-coding-agents.md#4-set-budgets">{limit, spent, estimate}</a>, and the whole session verifies cryptographically:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;markdown&quot;,&quot;nodeId&quot;:&quot;5bcf0ff7-5d79-4f50-b4e0-87cd68ac865e&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-markdown">  Session sess-coding-demo: 10 record(s), 10 valid, 0 invalid</code></pre></div><p>Leak 3 closed, with a receipt.</p><h2>Try it in 30 seconds (no API keys)</h2><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;bash&quot;,&quot;nodeId&quot;:&quot;2552a347-50ba-4d61-8a35-5d54174600f0&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-bash">git clone https://github.com/dativo-io/talon &amp;&amp; cd talon
make coding-agents-demo        # Docker; fully offline, deterministic</code></pre></div><p>The demo runs the whole story against a deterministic mock that speaks both providers&#8217; wire formats, including streaming and cache tokens. Every figure in this post can be recomputed from it. In a field where the loudest numbers are self-reported savings percentages and anonymous nine-figure invoice anecdotes, we&#8217;d rather hand you a demo than an assertion.</p><h3>What this doesn&#8217;t do</h3><p></p><ul><li><p style="text-align: justify;"><mark data-color="#d0e0e3" style="background-color: rgb(208, 224, 227); color: rgb(0, 0, 0);">Attribution is not authentication. </mark>Subagent names are client-asserted labels within an already-authenticated caller. Budgets bind to the caller and session &#8212; never to a label an agent could spoof.</p></li><li><p style="text-align: justify;"><mark data-color="#c9daf8" style="background-color: rgb(201, 218, 248); color: rgb(0, 0, 0);">The cap is soft.</mark> A single in-flight request can overshoot before the next one is denied. Hard reservation is on the roadmap.</p></li><li><p style="text-align: justify;"><mark data-color="#cfe2f3" style="background-color: rgb(207, 226, 243); color: rgb(0, 0, 0);">Local tools are invisible</mark>. <a href="https://dativo.io/talon/docs/configuration/">Talon governs what crosses it </a>- model and MCP traffic. Your agent&#8217;s shell commands run on the developer&#8217;s machine.</p></li></ul><h2>The takeaway</h2><p>Use <em>Fable 5</em> where a million tokens of context and frontier reasoning earn their price: the plan. </p><p>Use Opus and Codex where the cache makes tokens nearly free: the execution. That split is genuinely cost-effective &#8212; about $6.80 for a session that touches a 400k-token repo across 32 calls. But cost-<strong>effective</strong> only becomes cost-<strong>predictable</strong> when the session is a first-class thing: metered with cache-aware honesty, attributed per subagent across providers, and capped with a stop you can prove happened. That&#8217;s four HTTP headers and one line of YAML away. </p><p>*Talon is open source. The <a href="https://github.com/dativo-io/talon/blob/main/docs/guides/governing-coding-agents.md">coding-agents</a> guide has the full setup for Claude Code and Codex CLI;  the <a href="https://github.com/dativo-io/talon/tree/main/examples/coding-agents-demo">demo</a> leaves you holding the signed evidence.</p>]]></content:encoded></item><item><title><![CDATA[Frontier AI Is No Longer Globally Available by Default]]></title><description><![CDATA[A European 5 eurocents on a fragmented world, passport-gated compute, and why the best models are becoming a privilege.]]></description><link>https://blog.dativo.io/p/frontier-ai-is-no-longer-globally</link><guid isPermaLink="false">https://blog.dativo.io/p/frontier-ai-is-no-longer-globally</guid><dc:creator><![CDATA[Sergey]]></dc:creator><pubDate>Wed, 01 Jul 2026 12:30:07 GMT</pubDate><enclosure url="https://images.unsplash.com/photo-1751844107337-17a9648cbab0?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw3fHxubyUyMHRyZXNzcGFzaW5nfGVufDB8fHx8MTc4MjkwNzgwNnww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>For the last few years, the software engineering community treated frontier LLM endpoints like a commodity utility. You plug in an API key, send a prompt, and get a response. We bought into the myth of a flat, frictionless internet where state-of-the-art intelligence was uniformly distributed to whoever could pay the bill.</p><p>That era is over.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://images.unsplash.com/photo-1751844107337-17a9648cbab0?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw3fHxubyUyMHRyZXNzcGFzaW5nfGVufDB8fHx8MTc4MjkwNzgwNnww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://images.unsplash.com/photo-1751844107337-17a9648cbab0?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw3fHxubyUyMHRyZXNzcGFzaW5nfGVufDB8fHx8MTc4MjkwNzgwNnww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1751844107337-17a9648cbab0?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw3fHxubyUyMHRyZXNzcGFzaW5nfGVufDB8fHx8MTc4MjkwNzgwNnww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1751844107337-17a9648cbab0?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw3fHxubyUyMHRyZXNzcGFzaW5nfGVufDB8fHx8MTc4MjkwNzgwNnww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1751844107337-17a9648cbab0?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw3fHxubyUyMHRyZXNzcGFzaW5nfGVufDB8fHx8MTc4MjkwNzgwNnww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw"><img src="https://images.unsplash.com/photo-1751844107337-17a9648cbab0?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw3fHxubyUyMHRyZXNzcGFzaW5nfGVufDB8fHx8MTc4MjkwNzgwNnww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" width="3600" height="2400" data-attrs="{&quot;src&quot;:&quot;https://images.unsplash.com/photo-1751844107337-17a9648cbab0?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw3fHxubyUyMHRyZXNzcGFzaW5nfGVufDB8fHx8MTc4MjkwNzgwNnww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:2400,&quot;width&quot;:3600,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;U.s. property: no trespassing sign.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="U.s. property: no trespassing sign." title="U.s. property: no trespassing sign." srcset="https://images.unsplash.com/photo-1751844107337-17a9648cbab0?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw3fHxubyUyMHRyZXNzcGFzaW5nfGVufDB8fHx8MTc4MjkwNzgwNnww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1751844107337-17a9648cbab0?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw3fHxubyUyMHRyZXNzcGFzaW5nfGVufDB8fHx8MTc4MjkwNzgwNnww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1751844107337-17a9648cbab0?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw3fHxubyUyMHRyZXNzcGFzaW5nfGVufDB8fHx8MTc4MjkwNzgwNnww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1751844107337-17a9648cbab0?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw3fHxubyUyMHRyZXNzcGFzaW5nfGVufDB8fHx8MTc4MjkwNzgwNnww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Photo by <a href="https://unsplash.com/@dtrinksrph">David Trinks</a> on <a href="https://unsplash.com">Unsplash</a></figcaption></figure></div><p></p><p>We can stop debating <em>if</em> LLMs are an effective part of the software engineering stack. The business world has already bought into the velocity gains of autonomous coding engines and advanced completion systems. They are entrenched, and they are with us for the long haul.</p><p>But as these models transition into core production infrastructure, they are simultaneously fracturing along national borders. We are entering an era of <strong>passport-gated compute</strong>, where access to top-tier reasoning is no longer a standard web request, but a geopolitical privilege.</p><p>If your application architecture treats a specific foreign-hosted frontier model as a hardcoded, direct dependency<strong>, your systems are structurally vulnerable</strong> to a new kind of risk: systemic technological exclusion.</p><p></p><h2>The Threat of Structural Technological Exclusion</h2><p>When a model like Claude Fable 5 launches, it represents a fundamental shift away from simple chat windows toward long-horizon, autonomous execution. These architectures are engineered to compress months of multi-file codebase refactorings, complex migrations, and deep analytical compliance loops into a few hours of agentic execution.</p><p>Because the business side expects this <span>10 X times ( or what ever magic number)</span> development velocity as the new baseline, losing access to this class of compute is no longer a minor operational glitch.</p><p>The real danger in a fragmented world is not that an API endpoint drops for ten minutes while a cloud provider reboots a router. The danger is that non-US companies, international divisions, and foreign contractors face a sudden, legal exclusion from cutting-edge technology entirely.</p><p>If a US-based competitor can refactor their entire monolithic core in an afternoon using native access to Tier 1 frontier models, while your European or international team is restricted to running legacy, open-weights models locally due to export boundaries, you aren&#8217;t dealing with a downtime event. You are dealing with a permanent, compounding competitive penalty.</p><h2>Timeline of loosing it</h2><p>The rapid policy shifts surrounding recent model rollouts offer a stark demonstration of how quickly a flat developer ecosystem can dissolve under state pressure. The 18-day geopolitical stalemate between Anthropic and the U.S. government serves as a blueprint for the future:</p><ul><li><p><strong>June 12:</strong> Citing national security concerns over dual-use capabilities, the U.S. Commerce Department&#8217;s Bureau of Industry and Security issued an emergency export control directive under the <a href="https://labs.cloudsecurityalliance.org/research/csa-research-note-ai-model-export-controls-enterprise-govern/">Export Administration Regulations (EAR)</a>. It ordered Anthropic to immediately suspend access to Fable 5 and Mythos 5 for all foreign nationals&#8212;including Anthropic&#8217;s own overseas employees. Because user nationality cannot be verified dynamically at a raw API boundary, Anthropic was forced to implement a blanket global shutdown of the models to ensure compliance.</p></li><li><p><strong>The Shockwave:</strong> Engineering environments outside the U.S. vanished overnight. Teams spanning from Canada to Central Europe were instantly locked out of active development pipelines. A partial relaxation days later allowed access to be restored&#8212;but <em>only</em> for verified, trusted U.S. organizations via early deployment pipelines. The developer ecosystem was split into distinct tiers.</p></li><li><p><strong><span>July 1 (</span><a href="https://gemini.google.com/app/6bff438c7871accd#:~:text=credit%20paywall).-,Anthropic,-The%20Hyper%2DSensitive"><span>The Degraded Return</span></a><span>):</span></strong><span> While the U.S. Commerce Department officially lifted the export ban after intense weeks of coordination, the global &#8220;return&#8221; of Fable 5 proved that the era of frictionless compute is dead.</span> <span>Anthropic&#8217;s immediate deployment restrictions show the new normal: the model is heavily throttled (capped at 50% of limits through July 7 before hitting a strict usage-credit paywall).</span></p></li><li><p><strong><span>The Hyper-Sensitive Classifier:</span></strong><span> To satisfy state security mandates, Fable 5 now runs behind an incredibly </span><a href="https://thenewstack.io/how-anthropic-is-bringing-fable-5-back/"><span>aggressive defense-in-depth safety filter</span></a><span>.</span> <span>If an engineering prompt remotely mimics a potential security exploit or vulnerability edge case, the request is aggressively intercepted and downgraded.</span></p></li><li><p><strong>The Death of Zero-Data-Retention (ZDR):</strong> To feed these safety classifiers, Anthropic now mandates a strict 30-day input/output data retention window for both Fable 5 and Mythos 5. For enterprise teams dealing with proprietary codebases and rigid data sovereignty compliance, native access is effectively closed off by compliance design.</p></li></ul><p>The lesson for platform engineers is explicit: <strong>Frontier model access can change hourly based on compliance frameworks entirely outside your codebase, and identity validation at the passport level is the new baseline for compute.</strong></p><h2>Direct Model Access Is a Production Liability</h2><p>The standard integration pattern works perfectly for a local prototype:</p><blockquote><p>app &#8594; claude-fable-5</p></blockquote><p>In production, this pattern is brittle. If the upstream endpoint rejects an execution thread due to a sudden jurisdictional block, an identity token mismatch, or an updated export policy, your application pipeline drops dead.</p><p>Worse, makeshift fallback code written under pressure frequently creates catastrophic security vulnerabilities. If an automated engineering pipeline hits a geofence on Fable 5 and blindly reroutes the raw, proprietary payload to an alternative public endpoint that lacks explicit zero-data-retention (<strong>ZDR</strong>) terms, the system did not gracefully recover, it just traded an availability issue for an irreversible compliance and IP violation.</p><p>Hardcoding a single frontier model into an enterprise system is exactly like hardcoding a single cloud availability zone with no multi-region failover plan.</p><p><strong>The Brittle Reality of Silent Degradation:</strong> With 1st of July, 2026 rollout, if an autonomous engineering agent attempts a codebase-wide migration and its context payload triggers the hyper-sensitive safety classifier, the endpoint doesn&#8217;t just return a clean <code>403</code> error. It silently offloads the execution thread to the lower-tier Opus 4.8. For long-horizon agentic workflows engineered around Tier 1 reasoning, this silent capability drop causes pipelines to break mid-stream without a clear architectural failover.</p><h2>Request Capabilities, Not Model Names</h2><p>To survive a fragmented AI market, applications must be entirely decoupled from explicit model identifiers. Instead of hardcoding direct API calls, insert an internal abstraction layer:</p><blockquote><p>app &#8594; AI gateway &#8594; allowed model/provider</p></blockquote><p>The application layer should only request a functional <strong>capability</strong> and pass along the necessary operational context tokens:</p><ul><li><p><code>coding-frontier</code></p></li><li><p><code>legal-zdr</code></p></li><li><p><code>eu-sensitive</code></p></li><li><p><code>finance-deep-review</code></p></li></ul><p>The gateway resolves the optimal, compliant model instance at runtime by evaluating a real-time policy matrix:</p><ul><li><p><strong>Who</strong> is running the execution thread?</p></li><li><p><strong>Where</strong> is the tenant registered, and what is the data classification tier?</p></li><li><p><strong>What</strong> are the strict retention constraints of this specific payload?</p></li><li><p><strong>Is</strong> the primary target model currently accessible under these constraints?</p></li><li><p><strong>What</strong> is the exact, pre-approved fallback route if the primary target is blocked?</p></li></ul><p>The core architectural rule shifts from <em>&#8220;Send this to Fable 5&#8221;</em> to <strong>&#8220;Send this to the most capable model allowed for this specific payload context.&#8221;</strong></p><h2>Managing Context Without Identity Creep</h2><p>Consider a practical execution: An international development team based in Poland initiates a session with an internal coding agent. The app requests the <code>coding-frontier</code> capability. The gateway intercepts the payload and evaluates the environmental context:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;yaml&quot;,&quot;nodeId&quot;:&quot;dcc5dd8f-5201-4cc1-9a5f-f0f7d5e8bf0d&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-yaml">Context Evaluated:
- caller:            engineering-agent
- tenant:            EU company
- user jurisdiction: Poland
- data tier:         internal source code
- retention req:     zero data retention
- preferred model:   Claude Fable 5
- Fable status:      unavailable / restricted / no ZDR
</code></pre></div><p>The gateway instantly executes an explicit policy decision: <strong>Reject Fable 5.</strong> It automatically evaluates the approved fallback list, matching the payload to an available instance of Mistral, Azure OpenAI inside an EU region, or an enterprise Bedrock endpoint that strictly honors the zero-data-retention rule.</p><p>To do this safely, <strong>the gateway must not digest raw identity records.</strong> Shipping passport data or citizenship details inside an application payload is a severe anti-pattern. Instead, identity and HR platforms must map users to derived, lightweight entitlement flags before the gateway ever touches the request:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;yaml&quot;,&quot;nodeId&quot;:&quot;a80517a5-942f-4bf4-84ba-f85092e8c5a2&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-yaml">caller = engineering-agent
tenant = EU company
ai.fable_allowed = false
ai.eu_only = true
ai.zdr_required = true
data_tier = internal source code
</code></pre></div><p>The gateway does not need to know <em>why</em> <code>ai.fable_allowed</code> is false; it only needs to read the boolean token to route or deny the execution deterministically.</p><h2>Why the Gateway Must Be Internal Infrastructure</h2><p>When engineers realize they need an abstraction layer, the common temptation is to use a third-party cloud SaaS gateway. This is a severe architectural error that trades a configuration problem for a massive data-sovereignty vulnerability.</p><p>The core rule of distributed systems applies directly here: <strong>control planes can be remote, but data planes must be local.</strong></p><h3>1. The Compliance Paradox</h3><p>The underlying reason to build an AI routing layer is to guarantee regional data boundaries and enforce zero-data-retention rules. If you pass your traffic through a third-party cloud SaaS gateway, you are sending raw, unencrypted prompts, internal source code, and customer records to <em>another</em> intermediary entity before it reaches the final LLM provider. You have expanded your data processor surface area, not contained it.</p><h3>2. The Latency and Streaming Tax</h3><p>Agentic workflows live and die by Time to First Token (TTFT) metrics. Forcing your traffic through an external cloud proxy adds a redundant public internet hop, duplicate TCP handshakes, and extra TLS terminations. Furthermore, if that third-party SaaS buffers parts of the Server-Sent Events (SSE) stream to run out-of-band cost logging or analytics, interactive streaming performance drops immediately.</p><h3>3. Artificial Single Points of Failure</h3><p>An external SaaS gateway creates a mathematical availability bottleneck. Your total system uptime is no longer just dependent on your app and the final AI model - it is tied to the stability of the intermediate platform.</p><p>If the intermediate proxy experiences a routing loop or a regional cloud outage, your entire product fails&#8212;even if your core services and the underlying LLM endpoints are 100% healthy.</p><h3>4. Policy Locality and Context Hydration</h3><p>To make real-time routing choices, the gateway needs instant access to context. When hosted locally inside your own cluster or VPC perimeter, the gateway can query internal microservices or a local cache with sub-millisecond latency. An external cloud SaaS forces you to &#8220;hydrate&#8221; the payload, bundling sensitive tenant profiles and location metadata into outbound headers just so an external platform can parse them.</p><p>The correct architectural pattern is clear: run the gateway container as a local sidecar or an internal VPC ingress proxy. The data plane stays strictly within your perimeter, while an external control plane is used exclusively to pull down policy rulebooks and configuration updates asynchronously.</p><h2>Evidence and Auditability</h2><p>Every automated reroute or denial must be fully auditable. If your system silently switches models behind the scenes, you must be able to prove <em>why</em> that choice was made during post-incident reviews or compliance audits. The internal gateway should produce cryptographically signed telemetry records detailing:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;plaintext&quot;,&quot;nodeId&quot;:&quot;4a685edd-4692-4257-8e21-aeb3460d5656&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-plaintext">- requested capability &amp; requested model
- selected model &amp; rejected candidates
- explicit rejection reasons (e.g., policy_zdr_violation)
- caller identity &amp; tenant details
- policy engine version &amp; provider metadata state
- final decision status: allowed / denied / rerouted
</code></pre></div><p>Without this audit trail, you cannot answer fundamental operational questions: Why did our costs spike yesterday? Did sensitive source code leave our regional data boundary during a fallback event? Was our data processed in compliance with our customer service-level agreements?</p><h2>Building for the New Baseline</h2><p>This is precisely where specialized, self-hosted gateway engines like <a href="https://github.com/dativo-io/talon">Talon</a> fit into the modern enterprise stack.</p><p><a href="https://dativo.io/">Talon</a> deploys directly into your own infrastructure perimeter. It sits inline between your application code, autonomous agents, and upstream model providers. Before a single byte leaves your network, it identifies the caller, checks local policy, strips PII, enforces model allowlists, manages rate limits, strips unapproved agent tools, and signs the execution evidence.</p><p>Talon does not bypass government export controls or magically unlock restricted models for entities that are legally barred from using them. It solves the real engineering problem: ensuring your production applications do not break if a single provider endpoint alters its availability policies overnight. It turns frontier models from hardcoded, brittle dependencies into interchangeable, policy-validated infrastructure backends.</p><p>Advanced LLMs as coding assistants and autonomous engineering blocks are a permanent fixture of our industry. But because they are infrastructure, we can no longer afford to wire production code directly to them. Relying on the hope that an external API vendor&#8217;s regulatory landscape will remain globally uniform is not a design strategy.</p><p>Infrastructure requires routing, fallback patterns, policy enforcement, and auditable evidence. If frontier AI is now production infrastructure, it is time to start architecting like it.</p>]]></content:encoded></item><item><title><![CDATA[How to Manage AI Spending Before It Becomes Cloud Spend Again]]></title><description><![CDATA[Why AI Cost Control is Becoming an Infrastructure Problem]]></description><link>https://blog.dativo.io/p/how-to-manage-ai-spending-before</link><guid isPermaLink="false">https://blog.dativo.io/p/how-to-manage-ai-spending-before</guid><dc:creator><![CDATA[Sergey]]></dc:creator><pubDate>Tue, 30 Jun 2026 13:41:48 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!b80A!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8329036d-8a3a-4594-b148-ce1640a77c76_1024x608.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!b80A!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8329036d-8a3a-4594-b148-ce1640a77c76_1024x608.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!b80A!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8329036d-8a3a-4594-b148-ce1640a77c76_1024x608.png 424w, https://substackcdn.com/image/fetch/$s_!b80A!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8329036d-8a3a-4594-b148-ce1640a77c76_1024x608.png 848w, https://substackcdn.com/image/fetch/$s_!b80A!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8329036d-8a3a-4594-b148-ce1640a77c76_1024x608.png 1272w, https://substackcdn.com/image/fetch/$s_!b80A!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8329036d-8a3a-4594-b148-ce1640a77c76_1024x608.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!b80A!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8329036d-8a3a-4594-b148-ce1640a77c76_1024x608.png" width="1024" height="608" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8329036d-8a3a-4594-b148-ce1640a77c76_1024x608.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:608,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!b80A!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8329036d-8a3a-4594-b148-ce1640a77c76_1024x608.png 424w, https://substackcdn.com/image/fetch/$s_!b80A!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8329036d-8a3a-4594-b148-ce1640a77c76_1024x608.png 848w, https://substackcdn.com/image/fetch/$s_!b80A!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8329036d-8a3a-4594-b148-ce1640a77c76_1024x608.png 1272w, https://substackcdn.com/image/fetch/$s_!b80A!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8329036d-8a3a-4594-b148-ce1640a77c76_1024x608.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">AI Finops in a short</figcaption></figure></div><p>Over the last few months, I&#8217;ve seen the same problem appear across drastically different companies:</p><ul><li><p><strong>Startups</strong> trying to rein in spiraling coding-agent bills.</p></li><li><p><strong>SMBs</strong> aggressively adding AI to their support and operations pipelines.</p></li><li><p><strong>Enterprises</strong> opening the floodgates by giving employees access to multiple AI tools.</p></li></ul><p>Despite the different contexts, they are all asking the same question: <em>&#8220;How do we manage AI spending without blocking useful AI work?&#8221;</em></p><p>This question isn&#8217;t arising because AI failed. It&#8217;s happening because AI started working well enough to become expensive. As <a href="https://www.reuters.com/business/retail-consumer/cheaper-ai-is-better-soaring-bills-are-reshaping-how-businesses-choose-models-2026-06-29/">Reuters</a> recently noted, while base token prices are falling, real AI bills are still rising. Today&#8217;s tasks utilize longer contexts, more steps, massive files, multiple retries, and complex agentic workflows. Smart companies are beginning to shift routine work to cheaper models, reserving expensive frontier models for only the hardest tasks.</p><p>That is the right pattern. But to execute it, you need infrastructure.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://blog.dativo.io/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://blog.dativo.io/subscribe?"><span>Subscribe now</span></a></p><p></p><h2>Token Spend Behaves Exactly Like Cloud Spend</h2><p>AI spending is mirroring the early days of cloud computing. The cycle is predictable:</p><ol><li><p>Usage is heavily encouraged to drive innovation.</p></li><li><p>Every team adopts it.</p></li><li><p>A massive, unexpected bill arrives.</p></li><li><p>Finance scrambles to ask who owns the spend.</p></li></ol><p>The exact same phenomenon that happened with compute and storage is now happening with tokens. The <a href="https://www.wsj.com/cio-journal/how-companies-are-managing-ai-token-spend-833b6f7e">WSJ</a> recently highlighted companies applying traditional cloud FinOps patterns to AI token spend: dashboards, spend caps, showback, chargeback, smaller models, and strict usage accountability.</p><p>This is where the industry is heading. The solution is not to &#8220;ban AI&#8221; or blindly force the cheapest model everywhere. The solution is <strong>controlled AI usage</strong> categorized by team, application, model, and data type.</p><h3>The Core Problem: Direct Model Access</h3><p>During the experimentation phase, most AI usage looks like a spiderweb of direct connections:</p><ul><li><p><strong>Slack bot</strong> &#8594; OpenAI</p></li><li><p><strong>Coding agent</strong> &#8594; Anthropic</p></li><li><p><strong>Internal app</strong> &#8594; OpenAI</p></li><li><p><strong>Support tool</strong> &#8594; Claude</p></li><li><p><strong>Employee</strong> &#8594; Personal AI account</p></li><li><p><strong>Vendor AI</strong> &#8594; Unknown model</p></li></ul><p>This works well for prototypes, but it shatters in production. Every direct path creates its own siloed cost behavior, model preference, data retention assumption, API key, log system, and compliance risk. At a company scale, this architecture leaves leadership unable to answer basic questions:</p><ul><li><p>Which teams are spending the money?</p></li><li><p>Which models are they actually using?</p></li><li><p>Which requests contain PII?</p></li><li><p>Which vendor saw which sensitive data?</p></li></ul><p>This is why AI cost control cannot happen after the fact. It must happen at the traffic layer.</p><h2>The Practical Controls Companies Need</h2><p>If a company wants to manage AI spending properly, it needs to implement controls <em><strong>before</strong></em> the model call happens.</p><h3>1. Caller Identity</h3><p>Every AI request must be tied to an application or team identity (e.g., <code>support-slack-bot</code>, <code>engineering-coding-agent</code>, <code>finance-document-assistant</code>). Without caller identity, cost ownership is impossible.</p><h3>2. Per-Caller Budgets</h3><p>Each caller should have hard daily and monthly limits. This is the AI equivalent of cloud budget alerts, but the crucial difference is that the request is denied <em>before</em> the spend happens.</p><p>YAML</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;yaml&quot;,&quot;nodeId&quot;:&quot;7e1dd307-407f-4335-b3c6-dbc16c49073d&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-yaml">gateway:
  callers:
    - name: support-slack-bot
      tenant_id: support
      policy_overrides:
        max_daily_cost: 10.00
        max_monthly_cost: 200.00
</code></pre></div><h3>3. Model Allowlists</h3><p>Not every team needs access to a frontier model. Routine support, text classification, summarization, and internal search can operate perfectly on lighter models.</p><p>YAML</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;yaml&quot;,&quot;nodeId&quot;:&quot;86f2c576-e38c-40dc-a4a7-6b2ef18121d1&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-yaml">policy_overrides:
  allowed_models:
    - gpt-4o-mini
    - claude-haiku
</code></pre></div><h3>4. Intelligent Model Routing</h3><p>The golden rule is not &#8220;use the cheapest model.&#8221; The rule is: <strong>Use the cheapest model allowed for this specific request.</strong></p><ul><li><p>Public data can route to cheap models.</p></li><li><p>Sensitive data may require EU routing.</p></li><li><p>HR and legal tasks require strict zero-retention policies.</p></li></ul><h3>5. Context and Loop Limits</h3><p>Agentic coding and document workflows can burn tokens silently. Recent studies on agentic tasks show that token usage can vary by up to 30x for the exact same task&#8212;and higher usage doesn&#8217;t always equal higher accuracy. You need circuit breakers:</p><p>YAML</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;yaml&quot;,&quot;nodeId&quot;:&quot;ecffc3b3-5f27-4ef3-a007-0beca5329b8e&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-yaml">policies:
  resource_limits:
    max_iterations: 10
    max_tool_calls_per_run: 20
    max_cost_per_run: 0.50
</code></pre></div><h3>6. Semantic Caching</h3><p>If hundreds of users ask a similar question, the system should not call the LLM every time. Caching public or low-risk answers is the fastest way to slash inference costs. However, caching must be aware of context: strictly avoid caching PII, confidential data, or dynamic tool calls.</p><p></p><h2>Why Price Isn&#8217;t Enough: The Claude Fable 5 Example</h2><p>Claude Fable 5 perfectly illustrates why AI spending is fundamentally a routing problem, not just a cost problem.</p><p>Fable-class models are phenomenal for heavy lifting: deep analytics, massive context windows, and complex autonomous agents. However, they come with premium pricing and operational sensitivity. <a href="https://www.reuters.com/technology/microsoft-limits-employee-use-anthropics-claude-fable-5-over-data-retention-2026-06-10/?utm_source=chatgpt.com">Reuters</a> recently reported that companies like Microsoft have had to restrict employee use of advanced models like Fable 5 due to data-retention concerns.</p><p>The most capable model isn&#8217;t always the <em>right</em> model. While Fable 5 might be perfect for public code analysis, it may need to be blocked entirely for zero-data-retention workloads in HR or healthcare. This proves that model choice should never be hardcoded into an application; it must be dictated by a centralized policy.</p><h2>The Solution: Enter Talon</h2><p><a href="https://github.com/dativo-io/talon">Talon</a> introduces a much-needed control layer between your applications and the model providers. Instead of an unmanageable web of direct API calls, <a href="https://dativo.io/">Talon</a> centralizes the traffic.</p><p><strong>Before forwarding any request, Talon can automatically:</strong></p><ol><li><p>Identify the caller and evaluate their policy.</p></li><li><p>Check rate limits and estimate the cost.</p></li><li><p>Scan for PII and classify the data tier.</p></li><li><p>Filter dangerous tools and redact sensitive text.</p></li><li><p>Route the prompt to the cheapest allowed provider.</p></li><li><p>Record cryptographically signed evidence of the transaction.</p></li></ol><p>AI cost control should not be a depressing monthly report; it should be a real-time, automated runtime decision.</p><h2>What to Measure</h2><p>To effectively govern this new infrastructure, move away from looking at a single, terrifying AI bill. Instead, track these specific metrics to see exactly where your spend is coming from and whether your policies are working:</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!AiHT!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71d81dfc-4f13-4ff7-ae4e-af8da7cbe454_685x170.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!AiHT!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71d81dfc-4f13-4ff7-ae4e-af8da7cbe454_685x170.png 424w, https://substackcdn.com/image/fetch/$s_!AiHT!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71d81dfc-4f13-4ff7-ae4e-af8da7cbe454_685x170.png 848w, https://substackcdn.com/image/fetch/$s_!AiHT!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71d81dfc-4f13-4ff7-ae4e-af8da7cbe454_685x170.png 1272w, https://substackcdn.com/image/fetch/$s_!AiHT!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71d81dfc-4f13-4ff7-ae4e-af8da7cbe454_685x170.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!AiHT!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71d81dfc-4f13-4ff7-ae4e-af8da7cbe454_685x170.png" width="685" height="170" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/71d81dfc-4f13-4ff7-ae4e-af8da7cbe454_685x170.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:170,&quot;width&quot;:685,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:39710,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://blog.dativo.io/i/204275429?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71d81dfc-4f13-4ff7-ae4e-af8da7cbe454_685x170.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!AiHT!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71d81dfc-4f13-4ff7-ae4e-af8da7cbe454_685x170.png 424w, https://substackcdn.com/image/fetch/$s_!AiHT!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71d81dfc-4f13-4ff7-ae4e-af8da7cbe454_685x170.png 848w, https://substackcdn.com/image/fetch/$s_!AiHT!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71d81dfc-4f13-4ff7-ae4e-af8da7cbe454_685x170.png 1272w, https://substackcdn.com/image/fetch/$s_!AiHT!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71d81dfc-4f13-4ff7-ae4e-af8da7cbe454_685x170.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><h2>Final Thoughts</h2><p>Direct provider access is perfectly fine for early prototypes or low-risk, public-data applications. But once AI becomes shared, expensive, sensitive, or agentic, direct access is the wrong abstraction.</p><p>You cannot control AI spending by simply asking developers to &#8220;use fewer tokens&#8221;&#8212;that didn&#8217;t work for cloud compute, and it won&#8217;t work now.</p><p>Cheaper models, caching, and better prompting will help, but ultimately, companies need a centralized control point. AI is no longer just a neat tool employees play with; it is core traffic moving through your company&#8217;s nervous system. And traffic needs policy before it becomes spend.</p>]]></content:encoded></item><item><title><![CDATA[Your compliance records are missing your AI traffic]]></title><description><![CDATA[You have log entries for HR and CRM, but not the prompts going to OpenAI/Anthropic or Mistral. But your team may sends EU customer data to OpenAI every day. Where's the record?]]></description><link>https://blog.dativo.io/p/your-compliance-records-are-missing</link><guid isPermaLink="false">https://blog.dativo.io/p/your-compliance-records-are-missing</guid><dc:creator><![CDATA[Sergey]]></dc:creator><pubDate>Thu, 11 Jun 2026 21:18:26 GMT</pubDate><enclosure url="https://images.unsplash.com/photo-1762330467572-5199bc772a20?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2MHx8b3BlbmFpfGVufDB8fHx8MTc4MTA1NTA3OXww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>TL;DR</strong>: Every European / operating in EU company shall maintain a Record of Processing Activities. Almost none of them have an entry for the prompts their teams send to OpenAI,  Mistral or Deepseek every day &#8212; the fastest-growing processing activity in the building. You declare ten lines of YAML once in <a href="https://github.com/dativo-io/talon">Dativo Talon</a>; everything else is derived from records a consultant cannot fabricate. There&#8217;s a <a href="https://github.com/dativo-io/talon/tree/main/examples/auditor-pack">downloadable sample</a> pack you can hand to a reviewer today.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://images.unsplash.com/photo-1762330467572-5199bc772a20?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2MHx8b3BlbmFpfGVufDB8fHx8MTc4MTA1NTA3OXww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://images.unsplash.com/photo-1762330467572-5199bc772a20?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2MHx8b3BlbmFpfGVufDB8fHx8MTc4MTA1NTA3OXww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1762330467572-5199bc772a20?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2MHx8b3BlbmFpfGVufDB8fHx8MTc4MTA1NTA3OXww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1762330467572-5199bc772a20?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2MHx8b3BlbmFpfGVufDB8fHx8MTc4MTA1NTA3OXww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1762330467572-5199bc772a20?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2MHx8b3BlbmFpfGVufDB8fHx8MTc4MTA1NTA3OXww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw"><img src="https://images.unsplash.com/photo-1762330467572-5199bc772a20?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2MHx8b3BlbmFpfGVufDB8fHx8MTc4MTA1NTA3OXww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" width="6000" height="4000" data-attrs="{&quot;src&quot;:&quot;https://images.unsplash.com/photo-1762330467572-5199bc772a20?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2MHx8b3BlbmFpfGVufDB8fHx8MTc4MTA1NTA3OXww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:4000,&quot;width&quot;:6000,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Smartphone screen displays ai assistant options.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Smartphone screen displays ai assistant options." title="Smartphone screen displays ai assistant options." srcset="https://images.unsplash.com/photo-1762330467572-5199bc772a20?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2MHx8b3BlbmFpfGVufDB8fHx8MTc4MTA1NTA3OXww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1762330467572-5199bc772a20?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2MHx8b3BlbmFpfGVufDB8fHx8MTc4MTA1NTA3OXww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1762330467572-5199bc772a20?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2MHx8b3BlbmFpfGVufDB8fHx8MTc4MTA1NTA3OXww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1762330467572-5199bc772a20?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2MHx8b3BlbmFpfGVufDB8fHx8MTc4MTA1NTA3OXww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Photo by <a href="https://unsplash.com/@zulfugarkarimov">Zulfugar Karimov</a> on <a href="https://unsplash.com">Unsplash</a></figcaption></figure></div><p>A friend of mine runs platform engineering at a ~200-person B2B company in Germany. Earlier this year they were closing their biggest deal to date - enterprise customer, everything agreed except the security review.</p><p>Question 47 of the enterprise questionnaire shared in long excel: <em>&#8220;Describe how personal data in AI/LLM workflows is governed, including records of processing, sub-processors, and third-country transfers.&#8221;</em></p><p>He took it to their CTO. The CTO opened their information systems register, the register every European company is required to keep under GDPR Article 30( called there Record of Processing Activities - <a href="https://gdpr-info.eu/art-30-gdpr/">RoPA</a>) - it can be as simple as confluence, and found entries for the HR system, the CRM, the email marketing tool. Nothing about the support bot that had been summarizing customer tickets through AI for eight months. Oops.</p><p>&#8220;Where do our prompts go?&#8221; CTO asked. In theory , the knew that the support bot they purchased use OpenAI&#8217;s ChatGPT, but nobody in the room could answer with evidence where the prompts went and what was inside the prompts. There were logs, somewhere, spread across three SaaS dashboards. There was no record.</p><p>The deal closed five weeks late - five weeks of a platform engineering and a CTO reverse-engineering their own AI usage so they could write it down and have legal sign it. That&#8217;s the moment AI governance stops being a legal abstraction and becomes a sales blocker.</p><p></p><h2>Why AI traffic belongs in your Record of Processing Activities ( EU focused part)</h2><p>If you&#8217;re the engineer, focusing on EU, who has to answer question 47, here&#8217;s the 90-second version of what&#8217;s being asked.</p><p><strong>The Record of Processing Activities - RoPA (GDPR Art. 30)</strong> is a register that answers, per processing activity: what personal data do we process, why, about whom, who receives it, does it leave the EU, how long do we keep it, and how is it protected? It&#8217;s <em>mandatory</em> for nearly every European company. There&#8217;s a nominal under-250-employee exemption, but it doesn&#8217;t apply when processing is &#8220;not occasional&#8221; &#8212; and a support bot running every day is by definition not occasional. The RoPA is the first document requested in a regulator inquiry, an ISO 27001 surveillance audit, and most enterprise security reviews. Art. 30 violations sit in the fine tier of up to EUR 10M or 2% of global turnover.</p><p><strong>Annex IV (EU AI Act)</strong> is the technical documentation required for high-risk AI systems: what the system is, how it&#8217;s monitored and controlled, how risks are managed, how humans oversee it. High-risk obligations apply from <strong><mark data-color="#d9ead3" style="background-color: rgb(217, 234, 211); color: rgb(0, 0, 0);">August 2, 2026</mark></strong>. Most mid-size companies are <em>deployers</em> rather than providers, which means lighter obligations &#8212; but the documentation demand flows down through contracts anyway. Your enterprise customers will ask you to evidence your usage controls, oversight, and logging regardless of your formal role. Documentation-tier violations run up to EUR <em>15M or 3% of turnover</em>.</p><p>AI traffic is the gap in both documents. Companies have RoPA entries for systems built ten years ago, while prompts flowing to a US model provider &#8212; a new processing activity, with a new recipient, and often a third-country transfer &#8212; go unrecorded. It&#8217;s exactly the thing CTO/Legal/Compliance are now being asked about, and exactly the thing they can&#8217;t answer from scattered logs.</p><p>Here&#8217;s what I realized while building <a href="https://github.com/dativo-io/talon">Talon&#8217;s</a> gateway: <strong>the network layer already knows the answers.</strong> Which PII categories were observed in prompts. Which provider received them, in which region. What was redacted, what was blocked, what it cost. Talon records all of that per request, HMAC-signed at write time.</p><p>Some consultant writing your RoPA <em>guesses</em> at these facts. The gateway <em>proves</em> them.</p><p></p><h2>Declared + derived: ten lines of YAML, then evidence does the rest</h2><p><strong>Declared facts</strong> are business statements no log can know: who the controller is, why you process data, how long you keep it. Your DPO writes them once. Org-level identity goes in <code>talon.config.yaml</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;yaml&quot;,&quot;nodeId&quot;:&quot;55330662-24e6-4de2-a00d-2bba1c1bf6fa&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-yaml">compliance:
  controller:
    name: "Example GmbH"
    contact: "privacy@example.eu"
    dpo_contact: "dpo@example.eu"
    address: "Examplestr. 1, 10115 Berlin, Germany"</code></pre></div><p>Per-agent declarations live in <code>agent.talon.yaml</code>, next to the policy that governs the agent:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;yaml&quot;,&quot;nodeId&quot;:&quot;57ec4680-5af3-4655-996e-50960c82174f&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-yaml">compliance:
  frameworks: [gdpr, eu-ai-act]
  data_residency: eu
  declarations:
    processing:                      # GDPR Art. 30(1) facts
      purposes:
        - "customer support ticket triage"
      data_subject_categories:
        - "customers"
      personal_data_categories:
        - "contact details"
        - "payment identifiers"
        - "support ticket content"
      retention_period: "90 days after ticket closure"
      legal_basis: "contract (Art. 6(1)(b))"
      safeguards: "Role-based access; vendor DPAs on file; signed evidence retained for audit review"
    system:                          # EU AI Act Annex IV facts
      system_description: "Gateway-governed LLM assistant for support ticket triage"
      intended_purpose: "Summarize and route inbound support tickets"
      oversight_description: "Support lead reviews flagged tickets daily"</code></pre></div><p><strong>Derived facts</strong> come from the signed evidence store, and you never write them by hand: which processing activities actually ran (per tenant, per agent, first seen, last seen), which personal-data identifiers were actually observed, which recipients received data and in which region, which requests were third-country transfers, which policy denials fired, and which requests went through a human plan-review gate.</p><p>One design decision matters more than it looks: <strong>every governed request records where it went &#8212; not just the ones where PII was detected.</strong> A recipient list that depends on a classifier&#8217;s hit rate is a recipient list with silent holes; a missed identifier should never make a US provider disappear from your transfer table. Talon records the prompt &#8594; destination flow (provider, model, region) for all traffic &#8212; gateway requests, CLI and scheduled agent runs, MCP tool calls, and externally orchestrated graph runs &#8212; and layers the sensitivity classification on top. The rule <em>&#8220;a record claiming a model call must say where the data went&#8221;</em> is enforced as a runtime invariant on every evidence write, in CI by a parity contract test, and in the release smoke suite. The full controls-per-path breakdown is in the <a href="https://dativo.io/talon/docs/governance-control-matrix/">governance control matrix</a>.</p><p>Then one command merges declared and derived facts into each document:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;bash&quot;,&quot;nodeId&quot;:&quot;dc51c99c-bc14-41e0-a00d-bbfa5cf88c34&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-bash">talon compliance ropa --format html --output ropa.html
talon compliance annex-iv --format html --output annex-iv.html</code></pre></div><p>The HTML is print-to-PDF-ready. The JSON variant is machine-checkable, for the security reviewers who want to diff it between quarters.</p><h2>Missing declarations are flagged, not hidden</h2><h1>Your GDPR RoPA Is Missing Your AI Traffic &#8212; Here&#8217;s How to Fix It With Runtime Evidence</h1><p><strong>TL;DR:</strong> Every European company maintains a Record of Processing Activities. Almost none of them have an entry for the prompts their teams send to OpenAI every day &#8212; the fastest-growing processing activity in the building. <a href="https://github.com/dativo-io/talon">Dativo Talon</a>, an open-source AI governance gateway, now generates a GDPR Art. 30 RoPA and an EU AI Act Annex IV documentation pack directly from HMAC-signed runtime evidence: <code>talon compliance ropa</code> and <code>talon compliance annex-iv</code>. Your compliance officer declares ten lines of YAML once; everything else is derived from records a consultant cannot fabricate. There&#8217;s a <a href="https://github.com/dativo-io/talon/tree/main/examples/auditor-pack">downloadable sample pack</a> you can hand to a reviewer today.</p><div><hr></div><p>A friend of mine runs platform engineering at a ~400-person B2B company in Germany. Earlier this year they were closing their biggest deal to date &#8212; six figures, enterprise customer, everything agreed except the security review.</p><p>Question 47 of the questionnaire: <em>&#8220;Describe how personal data in AI/LLM workflows is governed, including records of processing, sub-processors, and third-country transfers.&#8221;</em></p><p>He took it to the compliance officer, who opened the company&#8217;s RoPA &#8212; the register every European company is required to keep under GDPR Article 30 &#8212; and found entries for the HR system, the CRM, the email marketing tool. Nothing about the support bot that had been summarizing customer tickets through GPT-4 for eight months.</p><p>&#8220;Where do our prompts go?&#8221; the compliance officer asked. Nobody in the room could answer with evidence. There were logs, somewhere, spread across three SaaS dashboards. There was no record.</p><p>The deal closed five weeks late &#8212; five weeks of a platform engineer and a compliance officer reverse-engineering their own AI usage so they could write it down and have legal sign it. That&#8217;s the moment AI governance stops being a legal abstraction and becomes a sales blocker.</p><h2>Why AI traffic belongs in your Record of Processing Activities</h2><p>If you&#8217;re the engineer who has to answer question 47, here&#8217;s the 90-second version of what&#8217;s being asked.</p><p><strong>The RoPA (GDPR Art. 30)</strong> is a register that answers, per processing activity: what personal data do we process, why, about whom, who receives it, does it leave the EU, how long do we keep it, and how is it protected? It&#8217;s mandatory for nearly every European company. There&#8217;s a nominal under-250-employee exemption, but it doesn&#8217;t apply when processing is &#8220;not occasional&#8221; &#8212; and a support bot running every day is by definition not occasional. The RoPA is the first document requested in a regulator inquiry, an ISO 27001 surveillance audit, and most enterprise security reviews. Art. 30 violations sit in the fine tier of up to EUR 10M or 2% of global turnover.</p><p><strong>Annex IV (EU AI Act)</strong> is the technical documentation required for high-risk AI systems: what the system is, how it&#8217;s monitored and controlled, how risks are managed, how humans oversee it. High-risk obligations apply from <strong>August 2, 2026</strong>. Most mid-size companies are <em>deployers</em> rather than providers, which means lighter obligations &#8212; but the documentation demand flows down through contracts anyway. Your enterprise customers will ask you to evidence your usage controls, oversight, and logging regardless of your formal role. Documentation-tier violations run up to EUR 15M or 3% of turnover.</p><p>AI traffic is the gap in both documents. Companies have RoPA entries for systems built ten years ago, while prompts flowing to a US model provider &#8212; a new processing activity, with a new recipient, and often a third-country transfer &#8212; go unrecorded. It&#8217;s exactly the thing compliance officers are now being asked about, and exactly the thing they can&#8217;t answer from scattered logs.</p><p>Here&#8217;s what I realized while building Talon&#8217;s gateway: <strong>the network layer already knows the answers.</strong> Which PII categories were observed in prompts. Which provider received them, in which region. What was redacted, what was blocked, what it cost. Talon records all of that per request, HMAC-signed at write time.</p><p>A consultant writing your RoPA <em>guesses</em> at these facts. The gateway <em>proves</em> them.</p><h2>Declared + derived: ten lines of YAML, then evidence does the rest</h2><p>Every auditor document splits into two kinds of facts, and Talon&#8217;s design keeps them strictly separate.</p><p><strong>Declared facts</strong> are business statements no log can know: who the controller is, why you process data, how long you keep it. Your compliance officer writes them once. Org-level identity goes in <code>talon.config.yaml</code>:</p><pre><code><code>compliance:
  controller:
    name: "Example GmbH"
    contact: "privacy@example.eu"
    dpo_contact: "dpo@example.eu"
    address: "Examplestr. 1, 10115 Berlin, Germany"</code></code></pre><p>Per-agent declarations live in <code>agent.talon.yaml</code>, next to the policy that governs the agent:</p><pre><code><code>compliance:
  frameworks: [gdpr, eu-ai-act]
  data_residency: eu
  declarations:
    processing:                      # GDPR Art. 30(1) facts
      purposes:
        - "customer support ticket triage"
      data_subject_categories:
        - "customers"
      personal_data_categories:
        - "contact details"
        - "payment identifiers"
        - "support ticket content"
      retention_period: "90 days after ticket closure"
      legal_basis: "contract (Art. 6(1)(b))"
      safeguards: "Role-based access; vendor DPAs on file; signed evidence retained for audit review"
    system:                          # EU AI Act Annex IV facts
      system_description: "Gateway-governed LLM assistant for support ticket triage"
      intended_purpose: "Summarize and route inbound support tickets"
      oversight_description: "Support lead reviews flagged tickets daily"</code></code></pre><p><strong>Derived facts</strong> come from the signed evidence store, and you never write them by hand: which processing activities actually ran (per tenant, per agent, first seen, last seen), which personal-data identifiers were actually observed, which recipients received data and in which region, which requests were third-country transfers, which policy denials fired, and which requests went through a human plan-review gate.</p><p>One design decision matters more than it looks: <strong>every governed request records where it went &#8212; not just the ones where PII was detected.</strong> A recipient list that depends on a classifier&#8217;s hit rate is a recipient list with silent holes; a missed identifier should never make a US provider disappear from your transfer table. Talon records the prompt &#8594; destination flow (provider, model, region) for all traffic &#8212; gateway requests, CLI and scheduled agent runs, MCP tool calls, and externally orchestrated graph runs &#8212; and layers the sensitivity classification on top. The rule <em>&#8220;a record claiming a model call must say where the data went&#8221;</em> is enforced as a runtime invariant on every evidence write, in CI by a parity contract test, and in the release smoke suite. The full controls-per-path breakdown is in the <a href="https://github.com/dativo-io/talon/blob/main/docs/reference/governance-control-matrix.md">governance control matrix</a>.</p><p>Then one command merges declared and derived facts into each document:</p><pre><code><code>talon compliance ropa --format html --output ropa.html
talon compliance annex-iv --format html --output annex-iv.html</code></code></pre><p>The HTML is print-to-PDF-ready. The JSON variant is machine-checkable, for the security reviewers who want to diff it between quarters.</p><h2>Missing declarations are flagged, not hidden</h2><p>This is the part I care most about as an engineer, and the part I would never bury in a demo screenshot.</p><p>If a declaration is missing, the command doesn&#8217;t fail &#8212; it renders a flagged <code>DECLARATION MISSING</code> section and prints exactly which YAML field to set. The document itself becomes the to-do list for your compliance officer. Talon fills what can be proven from signed evidence and clearly flags what must be declared by your organisation. That&#8217;s a far more trustworthy compliance story than &#8220;one-click compliance,&#8221; and it&#8217;s the kind of behaviour a compliance officer will actually trust after the first review.</p><p>The same discipline runs through the evidence-derived sections, in both directions:</p><ul><li><p><strong>No understatement.</strong> If no data-flow evidence exists yet, the transfers section says transfers <em>&#8220;cannot be assessed yet&#8221;</em> &#8212; it never converts absence of evidence into a comforting &#8220;no transfers&#8221; finding.</p></li><li><p><strong>No overstatement.</strong> When a policy blocks a request, the blocked attempt stays in the signed evidence, but the destination is <em>not</em> listed as a recipient &#8212; blocked data never reached anyone. A recipient table that counted blocked traffic would overstate your processing, and a compliance officer would catch it in the first review.</p></li><li><p><strong>Redaction is part of the record.</strong> An identifier type that was redacted in <em>every</em> flow to a destination is annotated <em>&#8220;redacted before egress&#8221;</em>; if it ever went through raw, even once, the annotation is withheld. &#8220;OpenAI received email addresses&#8221; and &#8220;OpenAI received placeholders where email addresses used to be&#8221; are very different statements to a compliance officer, and the document refuses to blur them.</p></li></ul><p>The document also checks your declarations against reality. If your agent declares <code>data_residency: eu</code> but routing still allows US providers and the evidence shows data actually flowing there, the RoPA prints a <strong>consistency warning</strong> with the two honest ways out: enforce <code>eu_strict</code> routing, or document the transfer mechanism with your compliance officer. Your own compliance export catches your config drift before an auditor does.</p><p>I watched all of this fire on a fresh install, unstaged. The first request I sent was denied at the routing stage &#8212; the agent&#8217;s policy pointed at a provider that wasn&#8217;t configured, so nothing ever left the machine. The denial landed in the signed evidence (<code>POLICY_DENIED_ROUTING</code>, with the fix spelled out in the record), but the RoPA generated from it listed <strong>no recipients</strong> and said transfers <em>&#8220;cannot be assessed yet&#8221;</em> &#8212; blocked data never reached anyone, so the document refused to invent a recipient. The second request succeeded against OpenAI, and the regenerated document did three things at once: put <code>openai / US</code> in the recipient table, flagged the third-country transfer with the SCC note, and opened with the consistency warning &#8212; because the config declared <code>data_residency: eu</code> while the evidence showed traffic reaching a US region:</p><blockquote><p><em>consistency: compliance.data_residency is declared &#8220;eu&#8221; but 1 destination(s) outside EU/LOCAL appear in data-flow evidence (Section 6) &#8212; set llm.routing.data_sovereignty_mode: eu_strict to enforce EU routing, or document the transfer mechanism (SCCs, adequacy decision) with your DPO</em></p></blockquote><p>That&#8217;s the document doing its job on run two of a brand-new install: catching a real residency gap I hadn&#8217;t noticed, and telling me exactly how to close it.</p><h2>The same evidence answers your security review, not just your compliance officer</h2><p>Question 47 rarely arrives alone. The same questionnaire that asks about records of processing asks how you prevent data leakage to AI providers, whether your audit logs can be tampered with, and what stops an AI agent from doing something it shouldn&#8217;t. Compliance documentation and infosec controls are usually owned by different people and answered from different tools &#8212; which is exactly why the answers contradict each other in review.</p><p>The reason a gateway can generate your RoPA is the same reason it closes the security gaps: it sits on the network path and enforces policy <em>before</em> data leaves, instead of reporting on it afterwards. Each control produces the same signed evidence the compliance documents are built from:</p><ul><li><p><strong>Shadow AI and unsanctioned usage.</strong> Talon gives AI traffic a single egress point. Every governed request is recorded with tenant, agent, destination, and region &#8212; whether or not PII was detected &#8212; so &#8220;what AI tools is the company actually using?&#8221; is a database query, not a survey. Traffic that bypasses the gateway is out of scope by definition, which is also your network team&#8217;s argument for routing it through.</p></li><li><p><strong>Data leakage to third parties.</strong> PII is detected and redacted <em>before</em> egress, and the evidence records both facts separately &#8212; input redaction and output redaction are independent claims. API keys never sit in prompts or agent code: they live in an encrypted vault (AES-256-GCM), scoped by per-agent ACLs, and every retrieval is itself an audit record.</p></li><li><p><strong>Prompt injection via attachments.</strong> File content (PDF, DOCX, HTML) is treated as untrusted by default &#8212; wrapped in isolation delimiters, scanned for embedded instructions, and blocked or flagged per policy. Detected injection attempts generate evidence even when the request is blocked.</p></li><li><p><strong>Excessive agency.</strong> Agents declare allowed tools in <code>agent.talon.yaml</code>; the policy engine filters the tool list before the model ever sees it, and every MCP tool call passes through policy evaluation before execution. A dangerous tool the model never learned about is a class of incident that can&#8217;t happen.</p></li><li><p><strong>Runaway spend.</strong> Per-request, daily, and monthly budgets are enforced at the gateway &#8212; a request over budget is denied before any cost is incurred, and the denial is recorded.</p></li><li><p><strong>Log tampering.</strong> Evidence records are HMAC-signed at write time. A reviewer &#8212; or your own incident-response team &#8212; can verify the export offline with <code>talon audit verify</code>. An attacker who modifies a record breaks its signature; &#8220;the logs say X and we can prove the logs weren&#8217;t altered&#8221; is a materially stronger incident report.</p></li></ul><p>This is also where the ISO 27001 and NIS2 story gets simple. The evidence store gives you logging and monitoring records (A.8.15), the vault covers cryptographic controls for credentials (A.8.24), the denial trail supports incident management (A.5.24&#8211;A.5.26), and NIS2 Art. 21&#8217;s risk-management measures get the same answer the compliance officer got: not a policy PDF, but signed records of the controls actually firing. <code>talon compliance report</code> prints the framework-to-control mappings alongside the runtime numbers, so your security lead and your compliance officer are quoting the same document for once.</p><h2>What manual AI compliance documentation actually costs</h2><p>If you&#8217;re the internal promoter &#8212; the platform engineer or CTO who has to convince a compliance officer, a board, or a customer&#8217;s security reviewer &#8212; here are the reference ranges (EU Commission impact assessment via CEPS, industry cost surveys, 2024&#8211;2026; all estimates):</p><ul><li><p><strong>Annex IV documentation done manually:</strong> EUR 15,000&#8211;60,000 per high-risk system; consultant-led packages run EUR 50,000&#8211;150,000+ over 3&#8211;6 months.</p></li><li><p><strong>DIY without tooling:</strong> EUR 30,000&#8211;80,000 in internal staff time &#8212; senior engineers writing documentation instead of product.</p></li><li><p><strong>RoPA maintenance:</strong> compliance-officer and privacy-consultant rates around EUR 800&#8211;1,500/day, days per review cycle, and the document goes stale the moment it&#8217;s written. Talon regenerates it from live evidence in one command.</p></li><li><p><strong>Compliance automation tools:</strong> EUR 7,500&#8211;80,000/yr &#8212; and none of them sit on the network path, so they produce templates, not evidence.</p></li></ul><p>But the biggest number isn&#8217;t on that list. For a 200&#8211;1,000-employee B2B company, the costliest compliance event isn&#8217;t a fine &#8212; it&#8217;s a <strong>stalled enterprise deal</strong>. &#8220;How is your AI usage governed?&#8221; is now a standard security-review question. Handing over a generated, independently verifiable evidence pack converts a five-week back-and-forth into an email attachment. One accelerated deal dwarfs every other line in this analysis.</p><p>The honest framing on fines: nobody can promise you &#8220;no fines,&#8221; and you should walk away from any vendor who does. What I can say is that when a regulator or auditor asks, the organization that produces organized, signed records in minutes is treated very differently from the one that produces nothing in weeks.</p><h2>What this does not do</h2><p>Claims discipline matters more in this domain than in any other:</p><ul><li><p>The output is <strong>supporting records for GDPR Art. 30 and EU AI Act Annex IV review</strong> &#8212; not a completed legal filing, a certification, or &#8220;compliance in a box.&#8221; Every document says so in its footer.</p></li><li><p>Talon only sees the traffic that flows through it. AI usage that bypasses the gateway isn&#8217;t in the evidence &#8212; the RoPA covers what Talon governs.</p></li><li><p>When an external agent framework (LangGraph, LangChain) is governed via Talon&#8217;s event API, the content itself never transits Talon &#8212; so those flows are recorded as exactly what they are: <em>orchestrator-reported</em>, model named, region <code>unknown</code>. Talon never guesses a jurisdiction. That unresolved region shows up in your transfer table on purpose.</p></li><li><p>You still need your compliance officer or counsel to review purposes, legal bases, and transfer mechanisms. The tool eliminates the <em>assembly</em> work &#8212; the weeks of reverse-engineering what your AI usage actually is &#8212; not the legal judgment.</p></li></ul><p>If a consultant told you a tool could replace them entirely, they&#8217;d be lying. If they told you assembling AI processing records by hand is a good use of EUR 1,200/day, they&#8217;d also be lying.</p><h2>Generate your first AI RoPA in ten minutes</h2><p>The whole point of Talon is that the proof path is short:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;bash&quot;,&quot;nodeId&quot;:&quot;90d969b1-a833-4992-bd13-bad799c2b3b7&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-bash"># 1. Init (wizard writes both config files, ~2 min)
talon init

# 2. Route a request through the gateway &#8212; this creates signed evidence
talon serve --gateway &amp;
curl -s http://localhost:8080/v1/proxy/openai/v1/chat/completions \
  -H "Authorization: Bearer $TALON_TENANT_KEY" \
  -d '{"model":"gpt-4o-mini","messages":[{"role":"user","content":"My IBAN is DE89370400440532013000"}]}'

curl -s http://localhost:8080/v1/proxy/openai/v1/chat/completions \
  -H "Authorization: Bearer $TALON_TENANT_KEY" \
  -d '{"model":"gpt-4o-mini","messages":[{"role":"user","content":"Hello, how are you?"}]}'


# 3. Generate the documents your compliance officer has been asked for
talon compliance report --format html --output compliance-report.html
talon compliance ropa --format html --output ropa.html
talon compliance annex-iv --format html --output annex-iv.html

# 4. Let the reviewer verify the evidence themselves
talon audit export --format signed-json --output evidence.signed.json
talon audit verify --file evidence.signed.json</code></pre></div><p>Open <code>ropa.html</code>. Section 4 already lists the IBAN identifier the classifier caught. Section 5 names the recipient and region &#8212; and if redaction was on, marks the IBAN <em>&#8220;redacted before egress&#8221;</em>: the provider got a placeholder, not the account number. Section 6 flags the third-country transfer with a note to document your SCC or adequacy mechanism. And if your declared residency disagrees with where the evidence shows traffic going, the document opens with the consistency warning from earlier.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!NnYc!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa60db9b6-71c2-47a9-aa42-5013ca7a508e_3364x1552.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!NnYc!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa60db9b6-71c2-47a9-aa42-5013ca7a508e_3364x1552.png 424w, https://substackcdn.com/image/fetch/$s_!NnYc!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa60db9b6-71c2-47a9-aa42-5013ca7a508e_3364x1552.png 848w, https://substackcdn.com/image/fetch/$s_!NnYc!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa60db9b6-71c2-47a9-aa42-5013ca7a508e_3364x1552.png 1272w, https://substackcdn.com/image/fetch/$s_!NnYc!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa60db9b6-71c2-47a9-aa42-5013ca7a508e_3364x1552.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!NnYc!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa60db9b6-71c2-47a9-aa42-5013ca7a508e_3364x1552.png" width="1456" height="672" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a60db9b6-71c2-47a9-aa42-5013ca7a508e_3364x1552.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:672,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:357537,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://blog.dativo.io/i/201630331?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa60db9b6-71c2-47a9-aa42-5013ca7a508e_3364x1552.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!NnYc!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa60db9b6-71c2-47a9-aa42-5013ca7a508e_3364x1552.png 424w, https://substackcdn.com/image/fetch/$s_!NnYc!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa60db9b6-71c2-47a9-aa42-5013ca7a508e_3364x1552.png 848w, https://substackcdn.com/image/fetch/$s_!NnYc!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa60db9b6-71c2-47a9-aa42-5013ca7a508e_3364x1552.png 1272w, https://substackcdn.com/image/fetch/$s_!NnYc!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa60db9b6-71c2-47a9-aa42-5013ca7a508e_3364x1552.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!iKw1!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3612a2bc-833f-478a-affe-0d6303011930_3394x1674.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!iKw1!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3612a2bc-833f-478a-affe-0d6303011930_3394x1674.png 424w, https://substackcdn.com/image/fetch/$s_!iKw1!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3612a2bc-833f-478a-affe-0d6303011930_3394x1674.png 848w, https://substackcdn.com/image/fetch/$s_!iKw1!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3612a2bc-833f-478a-affe-0d6303011930_3394x1674.png 1272w, https://substackcdn.com/image/fetch/$s_!iKw1!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3612a2bc-833f-478a-affe-0d6303011930_3394x1674.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!iKw1!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3612a2bc-833f-478a-affe-0d6303011930_3394x1674.png" width="1456" height="718" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3612a2bc-833f-478a-affe-0d6303011930_3394x1674.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:718,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:411579,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://blog.dativo.io/i/201630331?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3612a2bc-833f-478a-affe-0d6303011930_3394x1674.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!iKw1!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3612a2bc-833f-478a-affe-0d6303011930_3394x1674.png 424w, https://substackcdn.com/image/fetch/$s_!iKw1!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3612a2bc-833f-478a-affe-0d6303011930_3394x1674.png 848w, https://substackcdn.com/image/fetch/$s_!iKw1!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3612a2bc-833f-478a-affe-0d6303011930_3394x1674.png 1272w, https://substackcdn.com/image/fetch/$s_!iKw1!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3612a2bc-833f-478a-affe-0d6303011930_3394x1674.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p>If anyone asks about a single request, the answer is one record away &#8212; <code>talon audit show &lt;id&gt;</code> prints the per-request data flow in plain text:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;bash&quot;,&quot;nodeId&quot;:&quot;a1710738-e76f-428e-b26c-85cc79dc3eef&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-bash">PII Detected:  iban
PII Redacted:  input=true output=false
...
Data Flow
  Detector:    talon-regex
  prompt -&gt; llm_provider:openai model=gpt-4o region=US | redacted | tier 1 | iban</code></pre></div><p>Source, destination, region, what was detected, and whether it was redacted before it left &#8212; the RoPA&#8217;s recipient table, at the granularity of one request. Your compliance officer&#8217;s job goes from &#8220;reconstruct eight months of AI usage&#8221; to &#8220;review a document and fill in the fields the export flags.&#8221;</p><p>Don&#8217;t have the gateway running? Even the simplest path works: <code>talon run "hello"</code> against OpenAI is enough to put the provider in Section 5 and the US transfer in Section 6 &#8212; no PII required, because data movement is evidence regardless of what the data contained.</p><h2>FAQ</h2><p><strong>Does GDPR require a RoPA entry for AI and LLM usage?</strong><br>If your AI workflow processes personal data &#8212; support tickets, CRM notes, HR documents &#8212; it&#8217;s a processing activity under GDPR Art. 30 and belongs in your RoPA like any other system. The under-250-employee exemption doesn&#8217;t apply to processing that is &#8220;not occasional,&#8221; which rules out any AI feature running daily.</p><p><strong>Are prompts sent to OpenAI a third-country transfer?</strong><br>If the prompt contains personal data and the provider processes it outside the EU, yes &#8212; and your RoPA needs to record the recipient, region, and transfer mechanism (SCCs or an adequacy decision). This is exactly the section most companies cannot fill from scattered logs, and the one Talon derives from per-request data-flow evidence.</p><p><strong>Do SMBs need EU AI Act Annex IV documentation?</strong><br>Formally, Annex IV applies to providers of high-risk AI systems, and most SMBs are deployers with lighter obligations. In practice, enterprise customers push documentation demands down through contracts &#8212; you&#8217;ll be asked to evidence your usage controls, oversight, and logging in security reviews well before any regulator asks.</p><p><strong>Is a generated RoPA legally sufficient?</strong><br>It&#8217;s a supporting record, not a legal filing. Talon assembles the runtime facts (recipients, regions, identifiers observed, redaction status, denials) and flags the declarations only your organisation can make (legal basis, retention, purposes). Your compliance officer reviews and signs off &#8212; but starts from evidence instead of archaeology.</p><p><strong>Does this help with ISO 27001 and NIS2, or only GDPR?</strong><br>The same evidence store backs both. Signed per-request records support ISO 27001 logging and monitoring controls (A.8.15), the encrypted secrets vault maps to cryptographic controls (A.8.24), and the policy-denial trail supports incident management &#8212; which is also what NIS2 Art. 21 risk-management measures ask for. <code>talon compliance report</code> prints the framework-to-control mappings next to the runtime numbers.</p><p><strong>Where does the evidence live?</strong><br>In your infrastructure. Talon is a single Go binary, self-hosted and open source. Prompts, evidence records, and generated documents never leave your environment.</p><div><hr></div><ul><li><p>Repo: <a href="https://github.com/dativo-io/talon">github.com/dativo-io/talon</a> &#8212; single Go binary, self-hosted, open source</p></li><li><p>Runbook: <a href="https://dativo.io/talon/docs/compliance-export-runbook/">How to export evidence for auditors</a></p></li><li><p>Declarations guide: <a href="https://dativo.io/talon/docs/ropa-declarations/">How to clear DECLARATION MISSING blocks in RoPA exports</a></p></li><li><p><a href="https://gdpr-info.eu/art-30-gdpr/">GDPR Article 30 text</a> &#183; <a href="https://ai-act-service-desk.ec.europa.eu/en/ai-act/timeline/timeline-implementation-eu-ai-act">EU AI Act implementation timeline</a></p></li></ul><p>August 2, 2026 is on the calendar whether your RoPA is ready or not. The version of you that gets asked question 47 next quarter will be glad the answer is one command.</p>]]></content:encoded></item><item><title><![CDATA[Controlling LangGraph Tool Calls]]></title><description><![CDATA[Prompting Is Not Governance:]]></description><link>https://blog.dativo.io/p/controlling-langgraph-tool-calls</link><guid isPermaLink="false">https://blog.dativo.io/p/controlling-langgraph-tool-calls</guid><dc:creator><![CDATA[Sergey]]></dc:creator><pubDate>Tue, 02 Jun 2026 14:03:16 GMT</pubDate><enclosure url="https://images.unsplash.com/photo-1630783204535-cb30ffb3c0a1?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2MXx8dHJhZmZpYyUyMGxpZ2h0JTIwcG9saWNlbWVufGVufDB8fHx8MTc4MDQwNzIyMHww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Most AI agent demos stop precisely at the moment the agent &#8220;works.&#8221;</p><p>It can reason. It can choose tools. It can complete a linear workflow. For a proof of concept, this is a milestone. For enterprise production, it is barely the starting line.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://images.unsplash.com/photo-1630783204535-cb30ffb3c0a1?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2MXx8dHJhZmZpYyUyMGxpZ2h0JTIwcG9saWNlbWVufGVufDB8fHx8MTc4MDQwNzIyMHww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://images.unsplash.com/photo-1630783204535-cb30ffb3c0a1?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2MXx8dHJhZmZpYyUyMGxpZ2h0JTIwcG9saWNlbWVufGVufDB8fHx8MTc4MDQwNzIyMHww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1630783204535-cb30ffb3c0a1?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2MXx8dHJhZmZpYyUyMGxpZ2h0JTIwcG9saWNlbWVufGVufDB8fHx8MTc4MDQwNzIyMHww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1630783204535-cb30ffb3c0a1?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2MXx8dHJhZmZpYyUyMGxpZ2h0JTIwcG9saWNlbWVufGVufDB8fHx8MTc4MDQwNzIyMHww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1630783204535-cb30ffb3c0a1?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2MXx8dHJhZmZpYyUyMGxpZ2h0JTIwcG9saWNlbWVufGVufDB8fHx8MTc4MDQwNzIyMHww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw"><img src="https://images.unsplash.com/photo-1630783204535-cb30ffb3c0a1?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2MXx8dHJhZmZpYyUyMGxpZ2h0JTIwcG9saWNlbWVufGVufDB8fHx8MTc4MDQwNzIyMHww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" width="1080" height="607" data-attrs="{&quot;src&quot;:&quot;https://images.unsplash.com/photo-1630783204535-cb30ffb3c0a1?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2MXx8dHJhZmZpYyUyMGxpZ2h0JTIwcG9saWNlbWVufGVufDB8fHx8MTc4MDQwNzIyMHww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:607,&quot;width&quot;:1080,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;black traffic light with red light&quot;,&quot;title&quot;:&quot;black traffic light with red light&quot;,&quot;type&quot;:&quot;image/jpg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="black traffic light with red light" title="black traffic light with red light" srcset="https://images.unsplash.com/photo-1630783204535-cb30ffb3c0a1?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2MXx8dHJhZmZpYyUyMGxpZ2h0JTIwcG9saWNlbWVufGVufDB8fHx8MTc4MDQwNzIyMHww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1630783204535-cb30ffb3c0a1?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2MXx8dHJhZmZpYyUyMGxpZ2h0JTIwcG9saWNlbWVufGVufDB8fHx8MTc4MDQwNzIyMHww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1630783204535-cb30ffb3c0a1?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2MXx8dHJhZmZpYyUyMGxpZ2h0JTIwcG9saWNlbWVufGVufDB8fHx8MTc4MDQwNzIyMHww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1630783204535-cb30ffb3c0a1?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2MXx8dHJhZmZpYyUyMGxpZ2h0JTIwcG9saWNlbWVufGVufDB8fHx8MTc4MDQwNzIyMHww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Moving from prototype to production forces engineering teams to confront a fundamentally different architecture and risk question: <strong>What is this autonomous agent actually allowed to do?</strong></p><p>When an application switches from deterministic code to dynamic LLM runtime loops, traditional security models break down. A hallucinating chatbot can give a bad answer; an un-governed agent with access to tool arrays can inadvertently alter database states, leak source data, or trigger destructive cascade workflows.</p><p>To bridge this gap, engineers need to step away from fragile system prompts and build hard governance boundaries. In this walkthrough, we deploy a minimal <strong>LangGraph</strong> agent and introduce <strong><a href="https://github.com/dativo-io/talon">Talon</a></strong>&#8212;an OpenAI-compatible governance gateway&#8212;as a strict proxy between LangChain and the underlying LLM provider.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!oK6U!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa341bae2-c136-4c5a-8344-d0d966a4f64e_830x780.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!oK6U!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa341bae2-c136-4c5a-8344-d0d966a4f64e_830x780.png 424w, https://substackcdn.com/image/fetch/$s_!oK6U!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa341bae2-c136-4c5a-8344-d0d966a4f64e_830x780.png 848w, https://substackcdn.com/image/fetch/$s_!oK6U!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa341bae2-c136-4c5a-8344-d0d966a4f64e_830x780.png 1272w, https://substackcdn.com/image/fetch/$s_!oK6U!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa341bae2-c136-4c5a-8344-d0d966a4f64e_830x780.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!oK6U!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa341bae2-c136-4c5a-8344-d0d966a4f64e_830x780.png" width="830" height="780" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a341bae2-c136-4c5a-8344-d0d966a4f64e_830x780.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:780,&quot;width&quot;:830,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:105347,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://blog.dativo.io/i/200296967?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa341bae2-c136-4c5a-8344-d0d966a4f64e_830x780.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!oK6U!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa341bae2-c136-4c5a-8344-d0d966a4f64e_830x780.png 424w, https://substackcdn.com/image/fetch/$s_!oK6U!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa341bae2-c136-4c5a-8344-d0d966a4f64e_830x780.png 848w, https://substackcdn.com/image/fetch/$s_!oK6U!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa341bae2-c136-4c5a-8344-d0d966a4f64e_830x780.png 1272w, https://substackcdn.com/image/fetch/$s_!oK6U!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa341bae2-c136-4c5a-8344-d0d966a4f64e_830x780.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p>Our objective is to stress-test an essential architectural pattern: Can an infrastructure gateway inspect, filter, or block tool definitions downstream before the model ever encounters them?</p><p>TL;DR -  yes. And understanding why this is necessary requires exposing the core illusion of prompt-based guardrails.</p><h2>The Illusion of System Prompt &#8220;Enforcement&#8221;</h2><p>A pervasive anti-pattern in agent design is treating instructions as firewall rules. Developers routinely pass heavy system prompts down to the execution graph expecting deterministic obedience:</p><blockquote><p>You are a highly restricted enterprise assistant.</p><ul><li><p>Under no circumstances should you delete account records.</p></li><li><p>Do not expose or export sensitive PII or raw tables.</p></li><li><p>Always prompt the user for manual approval before mutating data.</p></li><li><p>Make no mistakes ;) LOL</p></li></ul></blockquote><p>While valuable for cognitive alignment, <strong>this is direction, not enforcement.</strong> If your backend payload still registers the schema definitions for <code>delete_record</code>, <code>export_data</code>, or <code>admin_override</code>, those tools are fully visible to the model context window. At that exact moment, your system security relies entirely on the probability that a non-deterministic token predictor will choose to follow instructions under every edge case, prompt injection vulnerability, or state variance.</p><div class="callout-block" data-callout="true"><p><strong>The Architectural Rule:</strong> If a tool schema is exposed to the model, the model can execute it. True governance dictates that restricted tools are stripped out at the gateway layer based on caller identity, making it physically impossible for the model to invoke what it cannot see.</p></div><p></p><h2>System Architecture &amp; Integration Setup</h2><p>Injecting a governance layer shouldn&#8217;t mean re-architecting your entire LangGraph state machine. By utilizing an OpenAI-compatible gateway like Talon, the code modifications are restricted to changing the initialization parameters of the LangChain client wrapper.</p><p>Instead of hitting the provider&#8217;s endpoint directly, we re-route traffic through our local or distributed proxy gateway:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;6ef46583-4f66-4fdc-a122-e98646b940e3&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from langchain_openai import ChatOpenAI

# Gateway-Routed LLM Client Configuration
llm = ChatOpenAI(
    model="gpt-4o-mini",
    base_url="http://localhost:18080/v1/proxy/openai/v1",  # Points to Talon Gateway
    api_key=TALON_CALLER_KEY,                             # Cryptographic Caller Identity Key
    temperature=0,
    max_tokens=120,
)</code></pre></div><h3>Key Security Mechanics:</h3><ol><li><p><strong>Abstraction of Secrets:</strong> The application container never handles the real <code>OPENAI_API_KEY</code>. It maintains a localized <code>TALON_CALLER_KEY</code>. The true downstream provider keys reside safely within Talon&#8217;s secure vault.</p></li><li><p><strong>Identity-Aware Routing:</strong> The gateway maps the inbound caller key to a specific tenant profile, resolves the associated governance policy, sanitizes the payload, and signs the outbound request to the provider.</p></li></ol><p></p><h2>Tool Inventories and the Gateway Policy</h2><p>To validate how the gateway behaves when handling complex real-world operations, our demonstration registers a blend of safe operational tools and highly sensitive data-access primitives:</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Qd7t!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3eab335-ad4b-42fb-9588-d666de3bf8aa_628x166.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Qd7t!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3eab335-ad4b-42fb-9588-d666de3bf8aa_628x166.png 424w, https://substackcdn.com/image/fetch/$s_!Qd7t!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3eab335-ad4b-42fb-9588-d666de3bf8aa_628x166.png 848w, https://substackcdn.com/image/fetch/$s_!Qd7t!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3eab335-ad4b-42fb-9588-d666de3bf8aa_628x166.png 1272w, https://substackcdn.com/image/fetch/$s_!Qd7t!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3eab335-ad4b-42fb-9588-d666de3bf8aa_628x166.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Qd7t!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3eab335-ad4b-42fb-9588-d666de3bf8aa_628x166.png" width="628" height="166" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d3eab335-ad4b-42fb-9588-d666de3bf8aa_628x166.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:166,&quot;width&quot;:628,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:40548,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://blog.dativo.io/i/200296967?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3eab335-ad4b-42fb-9588-d666de3bf8aa_628x166.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Qd7t!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3eab335-ad4b-42fb-9588-d666de3bf8aa_628x166.png 424w, https://substackcdn.com/image/fetch/$s_!Qd7t!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3eab335-ad4b-42fb-9588-d666de3bf8aa_628x166.png 848w, https://substackcdn.com/image/fetch/$s_!Qd7t!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3eab335-ad4b-42fb-9588-d666de3bf8aa_628x166.png 1272w, https://substackcdn.com/image/fetch/$s_!Qd7t!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3eab335-ad4b-42fb-9588-d666de3bf8aa_628x166.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p></p><p>When using standard LangGraph nodes, tools are bound directly via <code>.bind_tools(tools)</code>. LangChain automatically serializes these tool definitions into OpenAI-compliant JSON schemas.</p><p>Instead of relying on hardcoded static lists inside the code repository, we declare an external infrastructure policy in Talon (<code>policy.yaml</code>):</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;yaml&quot;,&quot;nodeId&quot;:&quot;fb422e77-ac77-4c2f-88d3-e862c9121ac5&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-yaml">callers:
  - name: "langgraph-tool-agent"
    tenant_key: "talon-gw-langgraph-tools-demo"
    tenant_id: "production-eu-west"
    allowed_providers:
      - "openai"
    policy_overrides:
      allowed_tools:
        - "search_records"
        - "update_record"
        - "send_notification"
      forbidden_tools:
        - "export_*"
        - "delete_*"
        - "admin_*"
        - "drop_*"
        - "truncate_*"
      allowed_models:
        - "gpt-4o-mini"</code></pre></div><h2>Why Pattern Matching Matters</h2><p>Relying on exact string matches for tools creates a brittle security stance. As engineering teams ship new features, developers might introduce variations like <code>delete_user</code>, <code>delete_workspace</code>, or <code>truncate_table</code>.</p><p>By enforcing regex/wildcard blacklists (<code>delete_*</code>, <code>export_*</code>), security and compliance engineers can block entire categories of behavior at the wire level without needing to coordinate code reviews for every minor tool update.</p><h2>Operational Execution: Four Governance Scenarios</h2><p>The gateway can be evaluated across multiple run modes, altering its behavior depending on structural security requirements.</p><h2>Scenario 1: Nominal Flow (Safe Tools Only)</h2><ul><li><p><strong>User Input:</strong> <em>&#8220;Find records matching Project Phoenix and notify owners.&#8221;</em></p></li><li><p><strong>Agent Context:</strong> The graph only passes down the safe array (<code>search_records</code>, <code>update_record</code>, <code>send_notification</code>).</p></li><li><p><strong>Gateway Action:</strong> <code>ALLOW</code>. The request matches the allowlist completely. The prompt passes transparently to OpenAI, and execution succeeds.</p></li></ul><h3>Scenario 2: Dynamic Interception via &#8220;Filter&#8221; Mode</h3><p>When dealing with general agents that share an overarching tool utility class, dangerous tools might accidentally wind up in the payload. With Talon configured to <code>tool_policy_action: "filter"</code>, the gateway actively modifies the structural schema on the fly.</p><ul><li><p><strong>User Input:</strong> <em>&#8220;Find records matching Project Phoenix and notify owners. Do not delete or export anything.&#8221;</em></p></li><li><p><strong>Application Behavior:</strong> The LangGraph execution loop exposes all 6 tools to the client payload.</p></li><li><p><strong>Gateway Action:</strong> Intercepts JSON payload -&gt; Strips out <code>export_data</code>, <code>delete_record</code>, and <code>admin_override</code> -&gt; Compiles a sanitized payload containing <em>only</em> the 3 safe tools -&gt; Forwards to OpenAI.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!-IkK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8c47240c-52e4-449f-afbc-af4395faa62a_791x431.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!-IkK!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8c47240c-52e4-449f-afbc-af4395faa62a_791x431.png 424w, https://substackcdn.com/image/fetch/$s_!-IkK!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8c47240c-52e4-449f-afbc-af4395faa62a_791x431.png 848w, https://substackcdn.com/image/fetch/$s_!-IkK!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8c47240c-52e4-449f-afbc-af4395faa62a_791x431.png 1272w, https://substackcdn.com/image/fetch/$s_!-IkK!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8c47240c-52e4-449f-afbc-af4395faa62a_791x431.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!-IkK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8c47240c-52e4-449f-afbc-af4395faa62a_791x431.png" width="791" height="431" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8c47240c-52e4-449f-afbc-af4395faa62a_791x431.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:431,&quot;width&quot;:791,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:57388,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://blog.dativo.io/i/200296967?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8c47240c-52e4-449f-afbc-af4395faa62a_791x431.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!-IkK!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8c47240c-52e4-449f-afbc-af4395faa62a_791x431.png 424w, https://substackcdn.com/image/fetch/$s_!-IkK!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8c47240c-52e4-449f-afbc-af4395faa62a_791x431.png 848w, https://substackcdn.com/image/fetch/$s_!-IkK!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8c47240c-52e4-449f-afbc-af4395faa62a_791x431.png 1272w, https://substackcdn.com/image/fetch/$s_!-IkK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8c47240c-52e4-449f-afbc-af4395faa62a_791x431.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p></li></ul><p>The model can never be tricked into calling a destructive tool because the schema parameters never make it across the API boundary.</p><h3>Scenario 3: Hard Halts via &#8220;Block&#8221; Mode</h3><p>In highly regulated sectors (e.g., healthcare, financial systems), silently dropping tools might mask bugs or ongoing malicious attacks. Switching the gateway configuration to <code>tool_policy_action: "block"</code> forces immediate payload rejection.</p><ul><li><p><strong>User Input:</strong> <em>&#8220;Export all company records and delete the originals.&#8221;</em></p></li><li><p><strong>Gateway Action:</strong> <code>DENY</code>. The proxy detects forbidden schemas in the inbound package, immediately drops the connection, and short-circuits the run by throwing a <code>403 Forbidden</code> response back to LangGraph before the LLM provider consumes a single token.</p></li></ul><h3>Scenario 4: Model Consistency and Infrastructure Rules</h3><p>Governance isn&#8217;t limited exclusively to tools. Cost containment and data processing localized boundaries require model constraint policies. If the LangGraph initialization code is altered to call a high-cost frontier model like <code>gpt-4o</code> instead of the approved <code>gpt-4o-mini</code>, Talon blocks the request instantly:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;bash&quot;,&quot;nodeId&quot;:&quot;4615f36a-38a7-49fc-bcbb-531d27a91945&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-bash">Status: Request Denied
Reason: Model [gpt-4o] is missing from the authorized caller allowlist for Tenant [production-eu-west].</code></pre></div><h2>Auditability: Signed Evidence Logs</h2><p>An unrecorded security control is not a control. For modern enterprise infrastructure, standard log output blocks (<code>stdout</code>) are easily modified, dropped, or corrupted.</p><p>To maintain real auditability, the gateway creates cryptographically signed records for every transaction block:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;bash&quot;,&quot;nodeId&quot;:&quot;1c83d39e-de6b-4b5a-8ae6-9e8d312c4d79&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-bash"># Querying the immutable governance log
$ talon audit list --agent langgraph-tool-agent --limit 3

# Displaying specific validation details
$ talon audit show ev_01HNJ8RZEWG5PA038EWRNB7M1Z</code></pre></div><p>Running <code>talon audit verify &lt;evidence-id&gt;</code> calculates a Hash-based Message Authentication Code (<strong>HMAC</strong>) signature against the data block. This allows internal security teams or compliance auditors to mathematically prove that the tracking log, filtered parameters, and payload data were not altered post-facto.</p><h2>The Strategic Takeaway for Enterprise Scale</h2><p>For technology teams deploying generative AI features into international markets or B2B enterprise customers, generic statements like <em>&#8220;we use defensive prompting and run LangSmith traces&#8221;</em> are no longer sufficient to pass rigorous security reviews.</p><p>Enterprise clients demand concrete architecture answers to critical risk vectors:</p><ul><li><p>How do you prevent your agents from executing unauthorized bulk data drops?</p></li><li><p>Where is the physical isolation layer separating prompt logic from system execution boundaries?</p></li><li><p>Where is the tamper-evident ledger tracking what your models attempted to execute?</p></li></ul><p>By decoupling <strong>orchestration</strong> from <strong>governance</strong>, you establish a resilient defense-in-depth security model:</p><ul><li><p><strong>LangGraph</strong> manages the state machine, execution graphs, memory persistence, and dynamic node routing.</p></li><li><p><strong><a href="https://dativo.io/">Talon</a> / API Gateways</strong> manage the network perimeter, secret isolation, tool schema sanitation, and cryptographic audit logging.</p></li></ul><p>This decoupling gives engineering teams the freedom to iterate rapidly on complex agent loops while giving security teams complete, granular control over the data boundaries. Prompting guides your agent&#8217;s behavior; policy enforces your system&#8217;s integrity.</p>]]></content:encoded></item><item><title><![CDATA[How to Make a LangGraph Agent GDPR-Safe]]></title><description><![CDATA[Use customer data in AI agents without losing control]]></description><link>https://blog.dativo.io/p/how-to-make-a-langgraph-agent-gdpr</link><guid isPermaLink="false">https://blog.dativo.io/p/how-to-make-a-langgraph-agent-gdpr</guid><dc:creator><![CDATA[Sergey]]></dc:creator><pubDate>Mon, 01 Jun 2026 21:02:43 GMT</pubDate><enclosure url="https://images.unsplash.com/photo-1606775524496-8ffd63ad2a98?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyNnx8ZXVyb3BlJTIwY29udHJvbHxlbnwwfHx8fDE3ODAzMzIzNTJ8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://images.unsplash.com/photo-1606775524496-8ffd63ad2a98?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyNnx8ZXVyb3BlJTIwY29udHJvbHxlbnwwfHx8fDE3ODAzMzIzNTJ8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://images.unsplash.com/photo-1606775524496-8ffd63ad2a98?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyNnx8ZXVyb3BlJTIwY29udHJvbHxlbnwwfHx8fDE3ODAzMzIzNTJ8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1606775524496-8ffd63ad2a98?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyNnx8ZXVyb3BlJTIwY29udHJvbHxlbnwwfHx8fDE3ODAzMzIzNTJ8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1606775524496-8ffd63ad2a98?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyNnx8ZXVyb3BlJTIwY29udHJvbHxlbnwwfHx8fDE3ODAzMzIzNTJ8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1606775524496-8ffd63ad2a98?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyNnx8ZXVyb3BlJTIwY29udHJvbHxlbnwwfHx8fDE3ODAzMzIzNTJ8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw"><img src="https://images.unsplash.com/photo-1606775524496-8ffd63ad2a98?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyNnx8ZXVyb3BlJTIwY29udHJvbHxlbnwwfHx8fDE3ODAzMzIzNTJ8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" width="2268" height="2442" data-attrs="{&quot;src&quot;:&quot;https://images.unsplash.com/photo-1606775524496-8ffd63ad2a98?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyNnx8ZXVyb3BlJTIwY29udHJvbHxlbnwwfHx8fDE3ODAzMzIzNTJ8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:2442,&quot;width&quot;:2268,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;man in blue crew neck shirt under blue sky during daytime&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="man in blue crew neck shirt under blue sky during daytime" title="man in blue crew neck shirt under blue sky during daytime" srcset="https://images.unsplash.com/photo-1606775524496-8ffd63ad2a98?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyNnx8ZXVyb3BlJTIwY29udHJvbHxlbnwwfHx8fDE3ODAzMzIzNTJ8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1606775524496-8ffd63ad2a98?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyNnx8ZXVyb3BlJTIwY29udHJvbHxlbnwwfHx8fDE3ODAzMzIzNTJ8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1606775524496-8ffd63ad2a98?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyNnx8ZXVyb3BlJTIwY29udHJvbHxlbnwwfHx8fDE3ODAzMzIzNTJ8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1606775524496-8ffd63ad2a98?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyNnx8ZXVyb3BlJTIwY29udHJvbHxlbnwwfHx8fDE3ODAzMzIzNTJ8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Photo by <a href="https://unsplash.com/@henri0019">Henri Lajarrige Lombard</a> on <a href="https://unsplash.com">Unsplash</a></figcaption></figure></div><p>This is a practical walkthrough for putting a minimal LangGraph agent behind the <a href="https://dativo.io/">Talon</a> LLM gateway.</p><p>The target use case is simple: a customer-support agent receives a billing question that contains personal data. We want the agent to answer, but we do not want the LangGraph app to call OpenAI directly with raw customer data and no audit trail.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://blog.dativo.io/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://blog.dativo.io/subscribe?"><span>Subscribe now</span></a></p><p></p><p>The final architecture:</p><pre><code>LangGraph &#8594; Talon Gateway &#8594; OpenAI</code></pre><p>The main application change is this:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;f9218104-0889-4d8a-81a7-8a2f17771061&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">llm = ChatOpenAI(
    model="gpt-4o-mini",
    base_url="http://localhost:18080/v1/proxy/openai/v1",
    api_key="talon-gw-langgraph-demo",
)</code></pre></div><p>The LangGraph app uses a Talon caller key. Talon stores the real OpenAI key, applies policy, forwards the request, scans the response, and writes signed evidence.</p><p>The companion notebook is here: <a href="https://github.com/dativo-io/talon-notebooks/blob/main/langgraph_talon_gdpr_safe_agent_colab.ipynb">Colab notebook</a></p><h2>What we are building</h2><p>We will run a minimal LangGraph workflow:</p><pre><code><code>START &#8594; support_agent &#8594; END</code></code></pre><p>No tools yet. No human approval. No memory.</p><p>That is intentional. The first thing to govern is the LLM boundary. Tool governance comes later.</p><p>The test input contains an email and an IBAN:</p><pre><code><code>My email is </code><strong>anna.kowalska@example.com</strong><code> and my IBAN is </code><strong>DE89370400440532013000</strong><code>.
I was charged twice for order </code><strong>ORD-18422</strong><code>. Can you help?</code></code></pre><p>What we want Talon to prove:<br><br>- Gateway receives the LangGraph request.<br>- Caller is identified as `langgraph-support-agent`.<br>- Model is restricted to `gpt-4o-mini`.<br>- Email address is detected.<br>- IBAN is detected.<br>- Input is redacted before the upstream provider call.<br>- Response is scanned.<br>- Evidence is written.<br>- Evidence signature verifies successfully.</p><div><hr></div><h2>Dependencies</h2><p>Python dependencies first:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;bash&quot;,&quot;nodeId&quot;:&quot;f878cc03-ff0d-46ad-9bf6-c286719f1c65&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-bash">python -m pip install -q --upgrade pip
python -m pip install -q langgraph langchain-openai langchain-core openai requests pyyaml</code></pre></div><p><a href="https://github.com/dativo-io/talon">Talon</a> also needs to be installed. One practical issue: do not use Ubuntu&#8217;s default <code>golang-go</code> package in Colab. It can be too old for current <a href="https://github.com/dativo-io/talon">Talon</a> builds.</p><p>The notebook tries three install paths:</p><ol><li><p>Use an existing <code>talon</code> binary if available.</p></li><li><p>Download a GitHub Linux AMD64 release binary.</p></li><li><p>Install modern Go from <code>go.dev</code> and run:</p></li></ol><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;bcf85f7f-bbaf-41cf-b8ff-11623c32f3e0&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">GOBIN=/usr/local/bin \
GONOSUMDB=github.com/dativo-io/talon \
go install github.com/dativo-io/talon/cmd/talon@latest</code></pre></div><p>That avoids the common failure where Colab installs Go 1.18 and the Talon module requires a newer Go version.</p><h2>Generate runtime keys</h2><p>For this demo, I generate ephemeral Talon keys inside the notebook:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;1768f99d-be26-49d0-89a2-a3540322cd7f&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">import os
import secrets
from pathlib import Path

project_dir = Path("/content/talon-langgraph-demo")
project_dir.mkdir(parents=True, exist_ok=True)

os.environ.setdefault("TALON_DATA_DIR", str(project_dir / ".talon"))
os.environ.setdefault("TALON_SECRETS_KEY", secrets.token_hex(32))
os.environ.setdefault("TALON_SIGNING_KEY", secrets.token_hex(32))
os.environ.setdefault("TALON_ADMIN_KEY", secrets.token_urlsafe(32))
os.environ.setdefault("TALON_PORT", "18080")</code></pre></div><p>For production, these should come from your secret manager. For a notebook, ephemeral keys are fine.</p><p>The OpenAI key is read from Colab Secrets if available:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;5809d185-135f-439f-ad95-0c6b03f2d2d8&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from google.colab import userdata

os.environ["OPENAI_API_KEY"] = userdata.get("OPENAI_API_KEY")</code></pre></div><p>or entered manually with <code>getpass</code>.</p><h2>Create <code>agent.talon.yaml</code></h2><p>This file describes the agent policy.</p><p>For this first example, keep it concise:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;yaml&quot;,&quot;nodeId&quot;:&quot;d7e43a72-951b-4ca9-8083-135704ef16d8&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-yaml">agent:
  name: langgraph-support-agent
  version: "1.0.0"
  description: GDPR-safe LangGraph support agent demo

capabilities:
  allowed_tools: []

policies:
  data_classification:
    input_scan: true
    output_scan: true
    redact_pii: true
    block_on_pii: false

  cost_limits:
    per_request: 0.10
    daily: 10.00
    monthly: 200.00

audit:
  log_level: detailed
  retention_days: 30
  log_prompts: false
  log_responses: false

compliance:
  frameworks:
    - gdpr
    - eu-ai-act
  data_residency: eu
  risk_level: low</code></pre></div><h2>Create <code>talon.config.yaml</code></h2><p>This file configures the gateway.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;yaml&quot;,&quot;nodeId&quot;:&quot;cfbbcec1-2dd9-4079-9b6a-7cc0f4f39885&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-yaml">gateway:
  mode: enforce

  providers:
    openai:
      enabled: true
      base_url: "https://api.openai.com"
      secret_name: "openai-api-key"
      allowed_models:
        - "gpt-4o-mini"

  default_policy:
    default_pii_action: "redact"
    response_pii_action: "warn"
    max_daily_cost: 10.00
    max_monthly_cost: 200.00
    allowed_models:
      - "gpt-4o-mini"

  callers:
    - name: "langgraph-support-agent"
      tenant_key: "talon-gw-langgraph-demo"
      tenant_id: "demo"
      allowed_providers:
        - "openai"
      policy_overrides:
        pii_action: "redact"
        response_pii_action: "warn"
        allowed_models:
          - "gpt-4o-mini"
        max_daily_cost: 10.00
        max_monthly_cost: 200.00</code></pre></div><p>The route we use later is:</p><pre><code><code>/v1/proxy/openai/v1/chat/completions</code></code></pre><p>Talon extracts <code>openai</code> from that path, looks it up under <code>gateway.providers.openai</code>, and checks that it is enabled.</p><h2>Store the OpenAI key in Talon&#8217;s vault</h2><p>The LangGraph app should not use the real OpenAI key.</p><p>Store the upstream provider key in Talon:</p><pre><code><code>talon secrets set openai-api-key "$OPENAI_API_KEY"</code></code></pre><p>Then the application uses the Talon caller key:</p><pre><code><code>talon-gw-langgraph-demo</code></code></pre><p>That key maps to:</p><pre><code><code>tenant_id: demo
name: langgraph-support-agent</code></code></pre><p>This gives you caller-level evidence and budget attribution.</p><h2>Start Talon Gateway</h2><p>In Colab, I use port <code>18080</code> to avoid collisions with common notebook services.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;bash&quot;,&quot;nodeId&quot;:&quot;3e0ea377-8581-4825-bf24-f2ade4fd104a&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-bash">talon serve \
  --gateway \
  --gateway-config /content/talon-langgraph-demo/talon.config.yaml \
  --host 127.0.0.1 \
  --port 18080 \
  --log-level info</code></pre></div><p>The OpenAI-compatible base URL becomes:</p><pre><code><code>http://localhost:18080/v1/proxy/openai/v1</code></code></pre><p>The full chat completions route is:</p><pre><code><code>http://localhost:18080/v1/proxy/openai/v1/chat/completions</code></code></pre><h2>Smoke-test the gateway before LangGraph</h2><p>Before involving LangGraph, test the gateway directly:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;4303c02a-4c11-4643-a4c0-dcba8c275ea8&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">import requests

base_url = "http://localhost:18080/v1/proxy/openai/v1"
url = f"{base_url}/chat/completions"

headers = {
    "Authorization": "Bearer talon-gw-langgraph-demo",
    "Content-Type": "application/json",
}

payload = {
    "model": "gpt-4o-mini",
    "messages": [
        {
            "role": "user",
            "content": (
                "My email is anna.kowalska@example.com and my IBAN is "
                "DE89370400440532013000. I was charged twice for order ORD-18422. "
                "Can you help?"
            ),
        }
    ],
    "max_tokens": 120,
}

r = requests.post(url, headers=headers, json=payload, timeout=60)
print(r.status_code)
print(r.text[:1500])</code></pre></div><p>Expected result: <code>200</code>.</p><p>Common failures:</p><p>- `404`: wrong route. Check `/v1/proxy/openai/v1/chat/completions`.</p><p>- `unknown or disabled provider`: missing `enabled: true` under `gateway.providers.openai`.</p><p>- `401` or `403`: wrong caller key, or missing admin key for admin endpoints.</p><p>- Upstream auth error: OpenAI key was not stored in Talon&#8217;s vault, or `secret_name` does not match.</p><p>- Schema validation error: invalid `agent.talon.yaml`; commonly `audit.log_level`.</p><div><hr></div><h2>Build the LangGraph agent</h2><p>The LangGraph agent is just one node:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;8ad1b94c-d097-4dec-9276-a458b138621f&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from typing import Annotated, TypedDict

from langchain_core.messages import HumanMessage, SystemMessage
from langchain_openai import ChatOpenAI
from langgraph.graph import StateGraph, START, END
from langgraph.graph.message import add_messages


class SupportState(TypedDict):
    messages: Annotated[list, add_messages]


llm = ChatOpenAI(
    model="gpt-4o-mini",
    base_url="http://localhost:18080/v1/proxy/openai/v1",
    api_key="talon-gw-langgraph-demo",
    temperature=0,
)


def support_agent(state: SupportState):
    system = SystemMessage(
        content=(
            "You are a customer support assistant for a SaaS company. "
            "Help with billing questions. "
            "Do not repeat raw personal data such as email addresses or IBANs. "
            "Do not claim you performed a refund. "
            "Say that a support teammate can verify the order and duplicate charge."
        )
    )

    response = llm.invoke([system] + state["messages"])
    return {"messages": [response]}


workflow = StateGraph(SupportState)
workflow.add_node("support_agent", support_agent)
workflow.add_edge(START, "support_agent")
workflow.add_edge("support_agent", END)

graph = workflow.compile()</code></pre></div><p>The Talon-specific part is only:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;d4b1c96d-667d-4159-8a6d-c96aaf54ecc8&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">base_url="http://localhost:18080/v1/proxy/openai/v1"
api_key="talon-gw-langgraph-demo"</code></pre></div><p>Everything else is standard LangGraph/LangChain code.</p><h2>Run the PII-bearing request</h2><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;0a0f1764-07b1-448c-a69e-74296964f55a&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">result = graph.invoke({
    "messages": [
        HumanMessage(
            content=(
                "My email is anna.kowalska@example.com and my IBAN is "
                "DE89370400440532013000. I was charged twice for order "
                "ORD-18422. Can you help?"
            )
        )
    ]
})

print(result["messages"][-1].content)</code></pre></div><p>At this point, the important thing is not whether the answer is amazing. This is a governance test, not a support automation benchmark.</p><p>The important question is: what did Talon record?</p><h2>Inspect evidence</h2><p>List records:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;bash&quot;,&quot;nodeId&quot;:&quot;af48d0a5-d8f4-4d45-a39e-244f0025e0bc&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-bash">talon audit list --limit 10</code></pre></div><p>You should see records with IDs like:</p><pre><code>gw_cd438803-4d0</code></pre><p>Then inspect one:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;bash&quot;,&quot;nodeId&quot;:&quot;17566287-340c-4e4b-b6be-392e05d48537&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-bash">talon audit show gw_cd438803-4d0</code></pre></div><p>A useful record should show:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;bash&quot;,&quot;nodeId&quot;:&quot;cd26cc28-0df0-4532-8782-e35ab6dbcb84&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-bash">Evidence:       gw_cd438803-4d0
Tenant / Agent: demo / langgraph-support-agent
Invocation:     gateway
HMAC Signature: &#10003; VALID

Policy Decision
Allowed:        true
Action:         allow

Classification
Input Tier:     2
Output Tier:    2
PII Detected:   email, iban, person, ...
PII Redacted:   true

Execution
Model:          gpt-4o-mini
Cost:           &#8364;&lt; 0.0001
Duration:       1267ms
Tokens:         in=106 out=62
Tools Called:   (none)</code></pre></div><p>This is the useful part of the demo.</p><p>The LLM call is no longer invisible. It has a tenant, an agent identity, a model, a PII classification, a redaction result, cost, latency, token counts, and a signature.</p><h2>Verify the evidence signature</h2><p>Run:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;bash&quot;,&quot;nodeId&quot;:&quot;5b5e8276-e676-41a7-923a-89c2d547e007&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-bash">talon audit verify gw_cd438803-4d0</code></pre></div><p>Expected:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;bash&quot;,&quot;nodeId&quot;:&quot;238f368d-06e8-4896-893f-66fb4e442fd1&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-bash">&#10003; Evidence gw_cd438803-4d0: signature VALID</code></pre></div><p>This proves the evidence record has not been modified since Talon created it.</p><p>For a technical buyer, this is materially different from application logs. Logs are useful for debugging. Signed evidence is useful for governance and later review.</p><h2>Understand the output PII warning</h2><p>In the audit explanation, you may see something like:</p><pre><code>POLICY_DENIED_PII_OUTPUT</code></pre><p>In this demo, that does not necessarily mean the request was blocked.</p><p>Check the final policy fields:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;yaml&quot;,&quot;nodeId&quot;:&quot;84b04522-5246-4005-903b-f06d3a9495ae&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-yaml">Allowed: true
Action: allow</code></pre></div><p>The reason is this config:</p><pre><code>response_pii_action: "warn"</code></pre><p>So Talon records output PII findings but still allows the response.</p><p>For the first demo, this is useful. It proves response scanning happened without making the notebook fail.</p><p>For production, pick the behavior explicitly:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;yaml&quot;,&quot;nodeId&quot;:&quot;267a075f-96cb-4dd1-b367-eedad6700002&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-yaml">response_pii_action: "warn"    # record only
response_pii_action: "redact"  # mask before returning
response_pii_action: "block"   # deny response</code></pre></div><p>One product note: the compact audit-list label can look confusing here. It would be clearer if the list view showed something like:</p><pre><code>ALLOWED_WITH_OUTPUT_PII_WARNING</code></pre><p>when the final action is allow but an output PII finding was recorded.</p><h2>Test model restriction</h2><p>The config only allows:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;yaml&quot;,&quot;nodeId&quot;:&quot;01f371f9-0875-4cab-a8b4-647245af16af&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-yaml">allowed_models:
  - "gpt-4o-mini"</code></pre></div><p>Test a different model:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;44acd1e2-aca8-4549-abee-1ae1fc2b9370&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">bad_payload = {
    "model": "gpt-4o",
    "messages": [{"role": "user", "content": "Say hello."}],
    "max_tokens": 20,
}

r = requests.post(url, headers=headers, json=bad_payload, timeout=60)
print(r.status_code)
print(r.text[:2000])</code></pre></div><p>Expected: a denial or policy error.</p><p>This is a basic but important control. You do not want every application instance choosing arbitrary models. Model selection affects cost, vendor review, latency, and sometimes data-residency posture.</p><h2>What this gives you</h2><p>This pattern gives a small team a practical governance boundary with a small app change.</p><p>What this adds:<br><br>- PII scan and redaction for raw customer data in prompts.<br>- Provider key isolation: the app calls <a href="https://dativo.io/">Talon</a>, not OpenAI directly.<br>- Model allowlist to prevent unapproved model usage.<br>- Caller identity mapped to tenant and agent.<br>- Signed evidence records for every governed call.<br>- Cost and token recording.<br>- Output scanning for response leakage risk.</p><div><hr></div><h2>Summary</h2><p>LangGraph makes it easy to add more autonomy: tools, loops, memory, retries, human approval, and long-running workflows.</p><p>Those are exactly the places where governance becomes harder.</p><p>Starting with the model boundary is the lowest-friction control:</p><pre><code><code>change base_url + api_key</code></code></pre><p>You can keep the LangGraph workflow mostly unchanged and still get:</p><pre><code><code>policy + redaction + model restriction + evidence</code></code></pre><p>That is the right first step before giving the agent tools that can read or write business data.</p><p>The next post will build on this and govern tool calls: safe tools, forbidden tools, dry runs, and blocking destructive actions before the agent can execute them.</p>]]></content:encoded></item><item><title><![CDATA[Data quality In Delta Lake and Iceberg]]></title><description><![CDATA[Part 3: Getting practical with data quality]]></description><link>https://blog.dativo.io/p/data-quality-in-delta-lake-and-iceberg-184</link><guid isPermaLink="false">https://blog.dativo.io/p/data-quality-in-delta-lake-and-iceberg-184</guid><dc:creator><![CDATA[Sergey]]></dc:creator><pubDate>Fri, 29 May 2026 06:44:48 GMT</pubDate><enclosure url="https://images.unsplash.com/photo-1568156318788-5c96955343a2?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw0MXx8aWNlYmVyZyUyMGxha2V8ZW58MHx8fHwxNzc5OTcyNjc1fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://images.unsplash.com/photo-1568156318788-5c96955343a2?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw0MXx8aWNlYmVyZyUyMGxha2V8ZW58MHx8fHwxNzc5OTcyNjc1fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://images.unsplash.com/photo-1568156318788-5c96955343a2?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw0MXx8aWNlYmVyZyUyMGxha2V8ZW58MHx8fHwxNzc5OTcyNjc1fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1568156318788-5c96955343a2?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw0MXx8aWNlYmVyZyUyMGxha2V8ZW58MHx8fHwxNzc5OTcyNjc1fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1568156318788-5c96955343a2?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw0MXx8aWNlYmVyZyUyMGxha2V8ZW58MHx8fHwxNzc5OTcyNjc1fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1568156318788-5c96955343a2?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw0MXx8aWNlYmVyZyUyMGxha2V8ZW58MHx8fHwxNzc5OTcyNjc1fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw"><img src="https://images.unsplash.com/photo-1568156318788-5c96955343a2?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw0MXx8aWNlYmVyZyUyMGxha2V8ZW58MHx8fHwxNzc5OTcyNjc1fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" width="3872" height="2160" data-attrs="{&quot;src&quot;:&quot;https://images.unsplash.com/photo-1568156318788-5c96955343a2?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw0MXx8aWNlYmVyZyUyMGxha2V8ZW58MHx8fHwxNzc5OTcyNjc1fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:2160,&quot;width&quot;:3872,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;gray mountains near pine trees at daytime&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="gray mountains near pine trees at daytime" title="gray mountains near pine trees at daytime" srcset="https://images.unsplash.com/photo-1568156318788-5c96955343a2?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw0MXx8aWNlYmVyZyUyMGxha2V8ZW58MHx8fHwxNzc5OTcyNjc1fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1568156318788-5c96955343a2?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw0MXx8aWNlYmVyZyUyMGxha2V8ZW58MHx8fHwxNzc5OTcyNjc1fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1568156318788-5c96955343a2?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw0MXx8aWNlYmVyZyUyMGxha2V8ZW58MHx8fHwxNzc5OTcyNjc1fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1568156318788-5c96955343a2?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw0MXx8aWNlYmVyZyUyMGxha2V8ZW58MHx8fHwxNzc5OTcyNjc1fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Photo by <a href="https://unsplash.com/@kenny_h">Kenneth Hargrave</a> on <a href="https://unsplash.com">Unsplash</a></figcaption></figure></div><p>This is the third part of series, the previous parts are - </p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;cc9bdcc1-6dc3-4870-a9df-7e8ece0b1d16&quot;,&quot;caption&quot;:&quot;Most companies already run some form of data quality monitoring. They have freshness checks, null checks, schema validation, row count checks, sometimes even anomaly detection, alerting, and incident workflows.&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;lg&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Data quality In Delta Lake and Iceberg&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:315636,&quot;name&quot;:&quot;Sergey&quot;,&quot;bio&quot;:&quot;Hey there! I'm Sergey Enin, a seasoned professional with 16+ years of experience in the advanced data analytics space. I've worked across the globe and I'm fluent in four languages &#128187;&#127757;&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cb4b8594-ff73-4b33-84a3-6ce5a2583e1c_144x144.png&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2026-05-28T12:27:00.008Z&quot;,&quot;cover_image&quot;:&quot;https://images.unsplash.com/photo-1605963476871-42dfd3bfc7af?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw3Mnx8aWNlYmVyZyUyMGZhbGxpbmd8ZW58MHx8fHwxNzc5OTY5MjYyfDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://blog.dativo.io/p/data-quality-in-delta-lake-and-iceberg&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:199590123,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:0,&quot;comment_count&quot;:0,&quot;publication_id&quot;:2240131,&quot;publication_name&quot;:&quot;Data, Engineering, and Beyond&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!SVdI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7bfbcde6-87b3-4b8c-9b38-3d1b82408e62_800x800.png&quot;,&quot;belowTheFold&quot;:false,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;140bb871-d3f3-45ec-be62-0abfbb8b2afa&quot;,&quot;caption&quot;:&quot;For data engineers, &#8220;data quality&#8221; is only one part of the operational picture.&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;lg&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Data quality In Delta Lake and Iceberg &quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:315636,&quot;name&quot;:&quot;Sergey&quot;,&quot;bio&quot;:&quot;Hey there! I'm Sergey Enin, a seasoned professional with 16+ years of experience in the advanced data analytics space. I've worked across the globe and I'm fluent in four languages &#128187;&#127757;&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cb4b8594-ff73-4b33-84a3-6ce5a2583e1c_144x144.png&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2026-05-28T12:59:23.112Z&quot;,&quot;cover_image&quot;:&quot;https://images.unsplash.com/photo-1713545690664-8dcba75d3c3e?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxfHxkZWx0YSUyMGxha2V8ZW58MHx8fHwxNzc5OTcxNDIyfDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://blog.dativo.io/p/data-quality-in-delta-lake-and-iceberg-e34&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:199594508,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:0,&quot;comment_count&quot;:0,&quot;publication_id&quot;:2240131,&quot;publication_name&quot;:&quot;Data, Engineering, and Beyond&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!SVdI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7bfbcde6-87b3-4b8c-9b38-3d1b82408e62_800x800.png&quot;,&quot;belowTheFold&quot;:false,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><p></p><p>Imagine a daily <code>finance.orders_mart</code> table used by executives.</p><p>The <a href="https://opendatacontract.com/">data contract</a> says:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;yaml&quot;,&quot;nodeId&quot;:&quot;19918f42-c3e0-4cf8-9cf2-ba97e7ca40ef&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-yaml">asset: finance.orders_mart
expected_schedule: daily
freshness_sla: available by 08:30 Europe/Warsaw
required_checks:
  - order_id_not_null
  - unique_order_id
  - revenue_non_negative
  - valid_currency
  - row_count_within_expected_range
owner: finance-data-platform
alert_route: "#data-finance-alerts"</code></pre></div><p>A pipeline run starts at 08:00.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://blog.dativo.io/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://blog.dativo.io/subscribe?"><span>Subscribe now</span></a></p><p></p><h2>Step 1: Check Upstream Freshness</h2><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;5caf9d8c-4e7c-46fe-b0c9-1ab652a8c11f&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">raw_orders_freshness = dq.get_freshness_status("raw.orders")

if raw_orders_freshness.status == "stale":
    stop_pipeline("raw.orders is stale")</code></pre></div><h2>Step 2: Run Transformation</h2><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;051ca4f2-4c82-49d7-9cc8-c382da987b09&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">orders = spark.table("raw.orders")
orders_mart = build_orders_mart(orders)

write_table(orders_mart, "finance.orders_mart")</code></pre></div><h2>Step 3: Check Pipeline Health</h2><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;1f7593f3-65ec-4876-85c3-fc1b0d6bfc46&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">pipeline_status = orchestrator.get_current_run_status()

if pipeline_status == "success":
    catalog.update_asset_metadata(
        asset="finance.orders_mart",
        pipeline_health="healthy",
        last_successful_pipeline_run_at=now()
    )</code></pre></div><h2>Step 4: Run DQ Checks</h2><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;b33a07ce-4a0e-46e5-a5f7-32fb0db535f4&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">dq_result = dq.run_checks(
    asset="finance.orders_mart",
    checks=[
        "order_id_not_null",
        "unique_order_id",
        "revenue_non_negative",
        "valid_currency",
        "row_count_within_expected_range"
    ]
)

dq_results_table.append(dq_result)</code></pre></div><h2>Step 5: Check Output Freshness</h2><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;7e3f3d3f-b5bc-4c6f-b1c3-2056abad4aef&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">freshness = dq.check_freshness(
    asset="finance.orders_mart",
    timestamp_column="order_created_at",
    expected_by="08:30",
    timezone="Europe/Warsaw"
)</code></pre></div><h2>Step 6: Publish Stable Asset State</h2><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;c97857cd-8567-430e-92e9-1b57b8cf22b0&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">if pipeline_status == "success" and dq_result.passed and freshness.passed:
    catalog.update_asset_metadata(
        asset="finance.orders_mart",
        data_contract_status="certified",
        quality_certification="certified",
        pipeline_health="healthy",
        freshness_state="fresh"
    )
else:
    catalog.update_asset_metadata(
        asset="finance.orders_mart",
        data_contract_status="warning"
    )</code></pre></div><p>The detailed operational results should stay in the right systems:</p><pre><code><strong>Pipeline logs </strong>     &#8594; orchestrator
<strong>DQ check results</strong>   &#8594; DQ platform / sidecar results table
<strong>Freshness history</strong>  &#8594; DQ platform / sidecar results table
<strong>Stable trust state</strong> &#8594; catalog / asset metadata</code></pre><p>The rule is simple:</p><pre><code><code>Operational evidence belongs in operational systems.
Stable trust state belongs in asset metadata.</code></code></pre><h2>The Combined Metadata Model</h2><p>For asset metadata, I suggest expose a compact view:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;yaml&quot;,&quot;nodeId&quot;:&quot;faa15718-4041-4df1-bbc0-e543bdb8b824&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-yaml">asset: finance.orders_mart

pipeline:
  health: healthy
  last_successful_run_at: 2026-05-27T08:12:00Z
  owner: finance-data-platform
  run_url: https://orchestrator/runs/123

quality:
  certification: certified
  contract_status: certified
  monitoring_required: true
  owner: finance-data-platform
  latest_results_uri: https://dq-platform/runs/456

freshness:
  state: fresh
  sla: available_by_08_30
  last_checked_at: 2026-05-27T08:20:00Z
  latest_results_uri: https://dq-platform/freshness/789</code></pre></div><p>This is enough for catalogs, BI tools, policies, and pipelines. It is not trying to store every check result.</p><h2>How Data Engineers Should Use Quality Indicators in Pipelines</h2><p>The most useful quality metadata is not decorative.</p><p>It should change how pipelines behave.</p><p>A tag like <code>quality_certification=certified</code> is only valuable if systems and people use it to make decisions. Otherwise, it is just another label in a catalog.</p><p>For data engineers, quality indicators can become pipeline control signals.</p><h2>1. Input Gating</h2><p>Before a pipeline reads from an upstream table, it can check the upstream asset state.</p><p>For example:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;001be3fe-e7ac-49aa-a943-bdb1eb3acd35&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">source_quality = catalog.get_asset_quality("raw.orders")

if source_quality.data_contract_status == "blocked":
    raise Exception("raw.orders is blocked by its data contract status")

if source_quality.quality_certification == "deprecated":
    warn("raw.orders is deprecated and should not be used for new pipelines")</code></pre></div><p>This does not require parsing every DQ result. The pipeline only needs a stable signal:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;yaml&quot;,&quot;nodeId&quot;:&quot;133fdbdd-bd5d-41e1-9a02-a7a310e70dc9&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-yaml">data_contract_status: certified | warning | blocked | deprecated</code></pre></div><p>This allows teams to prevent bad data from flowing silently into downstream tables.</p><h2>2. Freshness-Aware Execution</h2><p>Some pipelines should only run if the upstream data is fresh enough.</p><p>For example, a daily revenue table should not be rebuilt if the upstream orders table has not received today&#8217;s data.</p><p>The pipeline can check a freshness signal:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;yaml&quot;,&quot;nodeId&quot;:&quot;acd5e349-cf62-47b8-8f7d-9dadb544f0d7&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-yaml">freshness = dq_service.get_latest_check(
    asset="raw.orders",
    check="freshness"
)

if freshness.status == "failed":
    raise Exception("Upstream orders data is stale")</code></pre></div><p>There are two possible patterns here.</p><p>For critical operational freshness, read from the DQ platform directly because it has the latest check state.</p><p>For stable lifecycle decisions, read from the catalog or asset metadata.</p><p>That gives us a useful split:</p><pre><code><strong>Need latest operational status?</strong> &#8594; DQ platform
<strong>Need stable trust state?</strong>        &#8594; Catalog / asset metadata</code></pre><h2>3. Promotion Gates</h2><p>Quality indicators can/shall control movement between data layers.</p><p>For example:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;yaml&quot;,&quot;nodeId&quot;:&quot;c30bbd27-5660-44db-b632-45b15b22a087&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-yaml">bronze &#8594; silver
Requires:
- schema valid
- required fields present
- basic freshness check passing
- no severe ingestion errors

silver &#8594; gold
Requires:
- business rules passing
- accepted volume ranges
- key dimensions populated
- owner assigned

gold &#8594; certified
Requires:
- data contract approved
- monitoring enabled
- SLA defined
- alert routing configured
- successful check history</code></pre></div><p>This makes certification a real engineering workflow instead of a manual catalog label.</p><p>A pipeline could implement this as:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;dd48b609-1c01-4788-97c4-b1c9a7de712b&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">dq_result = dq.run_checks("curated.orders")

if dq_result.passed_required_checks:
    catalog.update_asset_metadata(
        asset="curated.orders",
        data_contract_status="managed"
    )

if dq_result.passed_certification_checks and owner_approved:
    catalog.update_asset_metadata(
        asset="curated.orders",
        quality_certification="certified"
    )</code></pre></div><p>The important part is that the pipeline does not write every check result into the table metadata.</p><p>It writes detailed results to the DQ system, then updates only the stable asset state when the lifecycle state changes.</p><h2>4. Output Validation</h2><p>Every important pipeline should validate what it produces.</p><p>This is where DQ tools are most useful.</p><p>After writing the output table, the pipeline runs checks such as:</p><pre><code><code>order_id is not null
revenue is non-negative
event_time is within expected freshness window
row count is within expected range
country_code matches known reference values
no duplicate primary business keys</code></code></pre><p>Then it publishes the result:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;528842da-e9f7-4278-b1a5-e019615a1ed0&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">dq_result = dq.run_checks(
    asset="finance.orders_mart",
    checks=[
        "order_id_not_null",
        "revenue_non_negative",
        "freshness_within_sla",
        "row_count_within_expected_range",
        "valid_country_code",
        "unique_order_id"
    ]
)

dq_results_table.append(dq_result)
dq_platform.publish(dq_result)</code></pre></div><p>The asset metadata may then be updated with a stable summary:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;yaml&quot;,&quot;nodeId&quot;:&quot;4960ae1f-7ecf-4848-8960-7311a36238b1&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-yaml">monitoring_required: true
quality_certification: certified
data_contract_status: certified
quality_owner: finance-platform</code></pre></div><p>But the run-level details stay outside the asset metadata.</p><h2>5. Incident-Aware Pipelines</h2><p>If a critical upstream asset has an open incident, downstream jobs can behave differently.</p><p>For example:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;yaml&quot;,&quot;nodeId&quot;:&quot;ad988b52-cd57-48b1-8ca6-65dd0caf6eca&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-yaml">incident = dq_service.get_open_incident("raw.payments")

if incident.severity == "critical":
    stop_pipeline()

if incident.severity == "warning":
    run_pipeline_but_mark_output_as_impacted()</code></pre></div><p>This enables a more nuanced model than simply &#8220;run or fail.&#8221;</p><p>Possible actions:</p><pre><code><strong>Critical failure</strong> &#8594; stop pipeline
<strong>Warning</strong>          &#8594; continue but mark output as impacted
<strong>Deprecated input </strong>&#8594; continue for existing jobs, block new dependencies
<strong>Freshness delay</strong>  &#8594; wait, retry, or skip publish
<strong>Schema break </strong>    &#8594; fail immediately</code></pre><p>This is how quality metadata becomes operationally useful.</p><h2>6. Lineage-Aware Impact Propagation</h2><p>The most powerful pattern is lineage-aware propagation.</p><p>If <code>raw.orders</code> fails, the platform should know that <code>curated.orders</code>, <code>finance.revenue_mart</code>, and <code>executive.arr_dashboard</code> may be impacted.</p><p>This does not mean all downstream tables immediately become &#8220;failed.&#8221; It means their trust state should reflect dependency risk.</p><p>For example:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;yaml&quot;,&quot;nodeId&quot;:&quot;7beab7c4-903d-4cdd-a8f1-dc84d78a7db7&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-yaml">quality_state: impacted
impacted_by: raw.orders
impact_reason: upstream_freshness_failed</code></pre></div><p>This is extremely useful for consumers.</p><p>Instead of discovering a broken dashboard manually, users can see that the underlying data is impacted by an upstream incident.</p><p>For data engineers, this also helps prioritize response. If a failed raw table impacts a board-level dashboard, it should be treated differently from a failure in an unused sandbox table.</p><h2>7. Environment and Release Gates</h2><p>Quality metadata can also be used in CI/CD workflows.</p><p>For example, a data contract change should not be promoted to production unless required checks exist.</p><p>A dbt model should not be marked as certified unless it has an owner, tests, freshness checks, and alert routing.</p><p>A table should not move from experimental to managed unless it has basic quality coverage.</p><p>This turns quality into a release policy:</p><pre><code><strong>No owner </strong><code>              &#8594; cannot certify
</code><strong>No freshness check</strong><code>     &#8594; cannot certify
</code><strong>No null checks</strong><code>         &#8594; cannot certify
</code><strong>No alert route </strong><code>        &#8594; cannot certify
</code><strong>Open critical incident</strong><code> &#8594; cannot promote</code></code></pre><p>This is where catalog metadata, DQ results, and pipeline orchestration come together.</p><h2>8. BI and Consumer Warnings</h2><p>Data engineers also need to think about downstream consumption.</p><p>A BI tool, notebook environment, or query interface can read asset trust metadata and display warnings.</p><p>For example:</p><pre><code>Warning<code>: this table is not certified.
Warning: this table is impacted by an upstream freshness incident.
Warning: this table is deprecated and will be removed after 2026-09-01.</code></code></pre><p>This is not a pipeline pattern, but data engineers need to publish the metadata that makes it possible.</p><h2>The Architecture I Would Recommend</h2><p>The clean architecture looks like this:</p><pre><code><code>Data pipeline
   &#8595;
Pipeline health checks
   &#8595;
DQ checks
   &#8595;
Freshness checks
   &#8595;
DQ platform / sidecar DQ results table
   &#8595;
Stable asset metadata update
   &#8595;
Catalog / governance layer
   &#8595;
Consumers, policies, BI warnings, certification workflows</code></code></pre><p>That is why it works.</p><h2>Final Take</h2><p>Data quality indicators should be part of the asset experience.</p><p>When someone opens a table, they should immediately understand whether it is certified, monitored, owned, trusted, stale, deprecated, blocked, or impacted by an upstream incident.</p><p>That information belongs in the catalog and can be mirrored into table metadata as stable properties.</p><p>At the same time, Data Quality operational results should not be embedded directly into Delta or Iceberg table metadata. They are too volatile, too detailed, and too operational. They can, and should, carry enough stable trust metadata to make data assets more discoverable, governable, and usable.</p><p></p><p></p><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://blog.dativo.io/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Data, Engineering, and Beyond! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Data quality In Delta Lake and Iceberg ]]></title><description><![CDATA[Part 2: Three Pillars of Data Engineering Monitoring]]></description><link>https://blog.dativo.io/p/data-quality-in-delta-lake-and-iceberg-e34</link><guid isPermaLink="false">https://blog.dativo.io/p/data-quality-in-delta-lake-and-iceberg-e34</guid><dc:creator><![CDATA[Sergey]]></dc:creator><pubDate>Thu, 28 May 2026 12:59:23 GMT</pubDate><enclosure url="https://images.unsplash.com/photo-1713545690664-8dcba75d3c3e?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxfHxkZWx0YSUyMGxha2V8ZW58MHx8fHwxNzc5OTcxNDIyfDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://images.unsplash.com/photo-1713545690664-8dcba75d3c3e?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxfHxkZWx0YSUyMGxha2V8ZW58MHx8fHwxNzc5OTcxNDIyfDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://images.unsplash.com/photo-1713545690664-8dcba75d3c3e?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxfHxkZWx0YSUyMGxha2V8ZW58MHx8fHwxNzc5OTcxNDIyfDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1713545690664-8dcba75d3c3e?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxfHxkZWx0YSUyMGxha2V8ZW58MHx8fHwxNzc5OTcxNDIyfDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1713545690664-8dcba75d3c3e?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxfHxkZWx0YSUyMGxha2V8ZW58MHx8fHwxNzc5OTcxNDIyfDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1713545690664-8dcba75d3c3e?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxfHxkZWx0YSUyMGxha2V8ZW58MHx8fHwxNzc5OTcxNDIyfDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw"><img src="https://images.unsplash.com/photo-1713545690664-8dcba75d3c3e?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxfHxkZWx0YSUyMGxha2V8ZW58MHx8fHwxNzc5OTcxNDIyfDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" width="4750" height="3160" data-attrs="{&quot;src&quot;:&quot;https://images.unsplash.com/photo-1713545690664-8dcba75d3c3e?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxfHxkZWx0YSUyMGxha2V8ZW58MHx8fHwxNzc5OTcxNDIyfDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:3160,&quot;width&quot;:4750,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;a lake surrounded by mountains under a cloudy sky&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="a lake surrounded by mountains under a cloudy sky" title="a lake surrounded by mountains under a cloudy sky" srcset="https://images.unsplash.com/photo-1713545690664-8dcba75d3c3e?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxfHxkZWx0YSUyMGxha2V8ZW58MHx8fHwxNzc5OTcxNDIyfDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1713545690664-8dcba75d3c3e?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxfHxkZWx0YSUyMGxha2V8ZW58MHx8fHwxNzc5OTcxNDIyfDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1713545690664-8dcba75d3c3e?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxfHxkZWx0YSUyMGxha2V8ZW58MHx8fHwxNzc5OTcxNDIyfDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1713545690664-8dcba75d3c3e?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxfHxkZWx0YSUyMGxha2V8ZW58MHx8fHwxNzc5OTcxNDIyfDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Photo by <a href="https://unsplash.com/@santurbanephotography">Abby Santurbane</a> on <a href="https://unsplash.com">Unsplash</a></figcaption></figure></div><p>For data engineers, &#8220;data quality&#8221; is only one part of the operational picture.</p><p>A production data asset is trustworthy only when three things are true:</p><p>1. <strong>Pipeline health</strong> - the pipeline which creates the data is healthy.</p><p>2. <strong>Data Quality </strong>- the data is correct enough for its use case.</p><p>3. <strong>Data Freshness</strong> - the data is fresh enough for its SLA.<br><br>They are not the same thing.</p><p>A pipeline can be green while the data is wrong.</p><p>A table can pass all schema and null checks while still being stale.</p><p>A freshness check can pass even if the pipeline is silently producing duplicated data.</p><p>Each pillar should produce operational signals, and only some of those signals should be promoted into stable asset metadata.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://blog.dativo.io/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://blog.dativo.io/subscribe?"><span>Subscribe now</span></a></p><p>This is the second part of series, the previous part is - </p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;6afd6c1f-d749-4844-b556-d3e71a460ae1&quot;,&quot;caption&quot;:&quot;Most companies already run some form of data quality monitoring. They have freshness checks, null checks, schema validation, row count checks, sometimes even anomaly detection, alerting, and incident workflows.&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;lg&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Data quality In Delta Lake and Iceberg&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:315636,&quot;name&quot;:&quot;Sergey&quot;,&quot;bio&quot;:&quot;Hey there! I'm Sergey Enin, a seasoned professional with 16+ years of experience in the advanced data analytics space. I've worked across the globe and I'm fluent in four languages &#128187;&#127757;&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cb4b8594-ff73-4b33-84a3-6ce5a2583e1c_144x144.png&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2026-05-28T12:27:00.008Z&quot;,&quot;cover_image&quot;:&quot;https://images.unsplash.com/photo-1605963476871-42dfd3bfc7af?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw3Mnx8aWNlYmVyZyUyMGZhbGxpbmd8ZW58MHx8fHwxNzc5OTY5MjYyfDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://blog.dativo.io/p/data-quality-in-delta-lake-and-iceberg&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:199590123,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:0,&quot;comment_count&quot;:0,&quot;publication_id&quot;:2240131,&quot;publication_name&quot;:&quot;Data, Engineering, and Beyond&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!SVdI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7bfbcde6-87b3-4b8c-9b38-3d1b82408e62_800x800.png&quot;,&quot;belowTheFold&quot;:true,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><h2>Pillar 1: Pipeline Health</h2><p>Pipeline health answers the question:</p><blockquote><p>Did the system that produces this asset run successfully?</p></blockquote><p>This is usually monitored by Airflow, dbt cloud, Spark jobs, Databricks Workflows, Flink, Kafka Connect, or another orchestration/runtime system.</p><p>Typical pipeline health signals:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;yaml&quot;,&quot;nodeId&quot;:&quot;f59cb11b-139a-4e94-8cde-add9390b8e19&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-yaml">pipeline.last_run_status: success | failed | skipped | running
pipeline.last_run_at: 2026-05-27T08:00:00Z
pipeline.last_successful_run_at: 2026-05-27T08:00:00Z
pipeline.duration_seconds: 842
pipeline.retries: 1
pipeline.owner: data-platform
pipeline.run_url: https://orchestrator/runs/123</code></pre></div><p>So, it could mean:</p><pre><code><code>The Airflow DAG finished successfully at 08:00.
The Spark job wrote the output table.
No task failed.
The pipeline SLA was met.</code></code></pre><p>This is good news, but it does not prove the data is correct.</p><p>A pipeline can complete successfully and still produce bad data because of upstream delays, incorrect joins, unexpected source changes, or silent business logic issues.</p><h3>Pipeline Health Check</h3><p>A data engineer may define a simple rule:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;cd6d8f55-bdd6-40e5-91df-6de4010aad2b&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">pipeline_run = orchestrator.get_latest_run("build_finance_orders_mart")

if pipeline_run.status != "success":
    catalog.update_asset_metadata(
        asset="finance.orders_mart",
        pipeline_health="failed",
        asset_state="not_trusted"
    )
    raise Exception("Pipeline failed")</code></pre></div><p>This signal is useful for operational dashboards and incident routing.</p><p>But I would not store every task log, retry event, executor metric, and stack trace in the table metadata. Those belong in the orchestrator and logging system.</p><p>The asset metadata should expose a compact state:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;yaml&quot;,&quot;nodeId&quot;:&quot;25de0ebc-b491-43c1-b7aa-43e94ae90edd&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-yaml">pipeline_health: healthy
last_successful_pipeline_run_at: 2026-05-27T08:00:00Z
pipeline_owner: finance-data-eng
pipeline_run_url: https://orchestrator/runs/123</code></pre></div><p></p><h2>Pillar 2: Data Quality</h2><p>Data quality answers the question:</p><blockquote><p>Did the produced data meet the expected rules?</p></blockquote><p>This is where tools such as Monte Carlo, Soda, Great Expectations, dbt tests, or even custom Spark checks are useful.</p><p>Typical data quality checks:</p><pre><code><code>Required columns are not null.
Primary business keys are unique.
Revenue is non-negative.
Country codes match reference data.
Order status belongs to an allowed list.
Row count is within an expected range.
Distribution of values did not unexpectedly shift.
Schema did not break.</code></code></pre><p>Example DQ result:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;yaml&quot;,&quot;nodeId&quot;:&quot;687e098d-7b28-4f97-bfdc-3bb3509f8c59&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-yaml">dq.status: failed
dq.failed_checks_count: 2
dq.failed_checks:
  - order_id_not_null
  - revenue_non_negative
dq.run_id: dq_run_123
dq.results_uri: https://dq-platform/runs/123</code></pre></div><p>This is operational state. It may change every time checks run.</p><p>The detailed result should stay in the data quality platform or sidecar DQ results table.</p><p>The asset metadata should expose only a stable summary or lifecycle signal:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;yaml&quot;,&quot;nodeId&quot;:&quot;95a696d3-ebec-4b6a-86d1-8b28a2a982de&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-yaml">quality_certification: certified
data_contract_status: warning
monitoring_required: true
quality_owner: finance-platform</code></pre></div><h3>Data Quality Gates</h3><p>A pipeline can run checks after producing a table:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;yaml&quot;,&quot;nodeId&quot;:&quot;1c2e4817-fbb9-4bb4-b50f-2398ff7dff39&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-yaml">orders_mart = build_orders_mart(raw_orders)

write_table(orders_mart, "finance.orders_mart")

dq_result = dq.run_checks(
    asset="finance.orders_mart",
    checks=[
        "order_id_not_null",
        "unique_order_id",
        "revenue_non_negative",
        "valid_order_status",
        "row_count_within_expected_range"
    ]
)

dq_results_table.append(dq_result)

if dq_result.has_critical_failures:
    catalog.update_asset_metadata(
        asset="finance.orders_mart",
        data_contract_status="blocked"
    )
    raise Exception("Critical DQ checks failed")

if dq_result.passed_required_checks:
    catalog.update_asset_metadata(
        asset="finance.orders_mart",
        data_contract_status="certified"
    )</code></pre></div><p><strong>The DQ run result is stored in the DQ results system , while certification status is stored in asset metadata.</strong></p><p>That keeps the asset metadata clean while still allowing pipelines to act on quality.</p><p></p><h2>Pillar 3: Data Freshness</h2><p>Data freshness answers the question:</p><blockquote><p>Is the data current enough for the business use case?</p></blockquote><p>Freshness is related to pipeline health, but it is not the same thing.</p><p>A pipeline may run successfully at 08:00, but if the upstream source stopped receiving new events at 02:00, the table is still stale.</p><p>Freshness checks usually compare expected arrival patterns with actual data timestamps or ingestion times.</p><p>Typical freshness signals:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;yaml&quot;,&quot;nodeId&quot;:&quot;97857649-1884-4798-9314-6346973f0665&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-yaml">freshness.status: fresh | stale | delayed | unknown
freshness.max_event_time: 2026-05-27T07:55:00Z
freshness.max_ingested_at: 2026-05-27T08:01:00Z
freshness.expected_by: 2026-05-27T08:15:00Z
freshness.delay_minutes: 20
freshness.sla_minutes: 60</code></pre></div><p>Example:</p><pre><code><code>The pipeline ran successfully.
The table has new rows.
But the newest business event is six hours old.
For a near-real-time reporting table, this is a freshness failure.</code></code></pre><h3>Freshness Gate</h3><p>For a critical dashboard table, a data engineer may define:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;e4548683-faf9-4baa-abca-6fed3c251d3f&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">freshness = dq.get_freshness_status("finance.orders_mart")

if freshness.delay_minutes &gt; freshness.sla_minutes:
    catalog.update_asset_metadata(
        asset="finance.orders_mart",
        freshness_state="stale",
        data_contract_status="warning"
    )
    notify_owner("finance.orders_mart is stale")</code></pre></div><p>For a downstream pipeline:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;c020e86d-1669-44ff-aa2e-4819080d5cc9&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">freshness = dq.get_freshness_status("raw.orders")

if freshness.status == "stale":
    raise Exception("Cannot rebuild finance.orders_mart because raw.orders is stale")</code></pre></div><p>Freshness is often the most important quality signal for business users.</p><p>A table can have perfect schema, zero nulls, and valid values, but if it is three days late, it is not useful.</p><h2>How the Three Pillars Work Together</h2><p>The real value comes from combining the three signals.</p><p>Consider this asset:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;yaml&quot;,&quot;nodeId&quot;:&quot;d52740d6-239c-48ff-aae7-c143f9a3c67a&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-yaml">asset: finance.orders_mart
pipeline_health: healthy
data_quality_status: passed
freshness_status: fresh
quality_certification: certified</code></pre></div><p>This table is in good shape.</p><div><hr></div><p></p><p>Now consider this one:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;yaml&quot;,&quot;nodeId&quot;:&quot;2516fdce-f0f4-49b7-bc62-cc484aa54f03&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-yaml">asset: finance.orders_mart
pipeline_health: healthy
data_quality_status: passed
freshness_status: stale
quality_certification: certified</code></pre></div><p>The pipeline is green and quality checks pass, but the data is stale. This should trigger a warning, especially for operational dashboards.</p><div><hr></div><p></p><p>Another case:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;yaml&quot;,&quot;nodeId&quot;:&quot;f04a1867-b2fa-427a-a138-2c0fd74e427c&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-yaml">asset: finance.orders_mart
pipeline_health: failed
data_quality_status: unknown
freshness_status: stale
quality_certification: certified</code></pre></div><p>The last run failed, so we do not know whether the latest data would pass checks. Freshness is stale because the table has not been updated. This is primarily a pipeline incident.</p><div><hr></div><p></p><p>Another case:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;yaml&quot;,&quot;nodeId&quot;:&quot;6c41fc85-0075-485b-860e-2e35108f379d&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-yaml">asset: finance.orders_mart
pipeline_health: healthy
data_quality_status: failed
freshness_status: fresh
quality_certification: suspended</code></pre></div><p>The pipeline ran and the data is fresh, but the values are wrong. This is a data quality incident.</p><div><hr></div><p></p><p>These distinctions matter because the response should be different.</p><blockquote><p><strong>Pipeline failed</strong>  &#8594; fix orchestration, runtime, permissions, infrastructure, code </p><p><strong>DQ failed  </strong>      &#8594; inspect data values, business rules, source changes, transformations </p><p><strong>Freshness failed </strong>&#8594; inspect source arrival, ingestion lag, scheduling, upstream dependencies</p></blockquote><p>This is why the metadata model should not simply say:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;yaml&quot;,&quot;nodeId&quot;:&quot;dca5a698-94d6-4c00-bf7c-12a93c1921c2&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-yaml">quality_status: failed</code></pre></div><p>That is too vague.</p><p>A better model says:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;yaml&quot;,&quot;nodeId&quot;:&quot;821add22-33c1-4687-8de5-887fa878e652&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-yaml">pipeline_health: healthy
data_quality_status: passed
freshness_status: stale
data_contract_status: warning</code></pre></div><p>Now we know what is wrong and triage it.</p>]]></content:encoded></item><item><title><![CDATA[Data quality In Delta Lake and Iceberg]]></title><description><![CDATA[Part 1: Stable Metadata + Operational Evidence]]></description><link>https://blog.dativo.io/p/data-quality-in-delta-lake-and-iceberg</link><guid isPermaLink="false">https://blog.dativo.io/p/data-quality-in-delta-lake-and-iceberg</guid><dc:creator><![CDATA[Sergey]]></dc:creator><pubDate>Thu, 28 May 2026 12:27:00 GMT</pubDate><enclosure url="https://images.unsplash.com/photo-1605963476871-42dfd3bfc7af?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw3Mnx8aWNlYmVyZyUyMGZhbGxpbmd8ZW58MHx8fHwxNzc5OTY5MjYyfDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://images.unsplash.com/photo-1605963476871-42dfd3bfc7af?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw3Mnx8aWNlYmVyZyUyMGZhbGxpbmd8ZW58MHx8fHwxNzc5OTY5MjYyfDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://images.unsplash.com/photo-1605963476871-42dfd3bfc7af?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw3Mnx8aWNlYmVyZyUyMGZhbGxpbmd8ZW58MHx8fHwxNzc5OTY5MjYyfDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1605963476871-42dfd3bfc7af?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw3Mnx8aWNlYmVyZyUyMGZhbGxpbmd8ZW58MHx8fHwxNzc5OTY5MjYyfDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1605963476871-42dfd3bfc7af?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw3Mnx8aWNlYmVyZyUyMGZhbGxpbmd8ZW58MHx8fHwxNzc5OTY5MjYyfDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1605963476871-42dfd3bfc7af?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw3Mnx8aWNlYmVyZyUyMGZhbGxpbmd8ZW58MHx8fHwxNzc5OTY5MjYyfDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw"><img src="https://images.unsplash.com/photo-1605963476871-42dfd3bfc7af?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw3Mnx8aWNlYmVyZyUyMGZhbGxpbmd8ZW58MHx8fHwxNzc5OTY5MjYyfDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" width="3008" height="2000" data-attrs="{&quot;src&quot;:&quot;https://images.unsplash.com/photo-1605963476871-42dfd3bfc7af?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw3Mnx8aWNlYmVyZyUyMGZhbGxpbmd8ZW58MHx8fHwxNzc5OTY5MjYyfDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:2000,&quot;width&quot;:3008,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;white ice on body of water&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="white ice on body of water" title="white ice on body of water" srcset="https://images.unsplash.com/photo-1605963476871-42dfd3bfc7af?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw3Mnx8aWNlYmVyZyUyMGZhbGxpbmd8ZW58MHx8fHwxNzc5OTY5MjYyfDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1605963476871-42dfd3bfc7af?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw3Mnx8aWNlYmVyZyUyMGZhbGxpbmd8ZW58MHx8fHwxNzc5OTY5MjYyfDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1605963476871-42dfd3bfc7af?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw3Mnx8aWNlYmVyZyUyMGZhbGxpbmd8ZW58MHx8fHwxNzc5OTY5MjYyfDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1605963476871-42dfd3bfc7af?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw3Mnx8aWNlYmVyZyUyMGZhbGxpbmd8ZW58MHx8fHwxNzc5OTY5MjYyfDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Photo by <a href="https://unsplash.com/@amelia1">Claudia Salvioli</a> on <a href="https://unsplash.com">Unsplash</a></figcaption></figure></div><p>Most companies already run some form of data quality monitoring. They have freshness checks, null checks, schema validation, row count checks, sometimes even anomaly detection, alerting, and incident workflows.</p><p>The problem is not that quality signals do not exist.</p><p>The problem is that they are usually hidden in operational tools, disconnected from the data assets people actually consume. So, data quality is a vanity metrics which is represented somewhere in the confluence and proudly shown to the leadership. But it means nothing.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://blog.dativo.io/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://blog.dativo.io/subscribe?"><span>Subscribe now</span></a></p><p></p><p>A data analyst opens a table in a catalog and sees the owner(finger crossed) , description, lineage, maybe some tags. But the real question is usually much simpler:</p><blockquote><p>Can I trust this table?</p></blockquote><p>That question leads to a idea:</p><p>Should data quality indicators become part of the table metadata itself?</p><p>Should Delta Lake or Apache Iceberg expose quality status directly as part of the asset?</p><p>After looking at this from the perspective of open table formats, catalogs, data quality platforms, and data engineering workflows, my conclusion is:</p><p><strong>Yes, data quality should be visible as asset metadata. But no, run-by-run data quality results should not be embedded directly into the open table format.</strong></p><p>The distinction matters.</p><h1>The Two Kinds of Data Quality Metadata</h1><p></p><p>When people say &#8220;data quality metadata,&#8221; they often mix two very different things.</p><p>The first category is stable asset state:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;yaml&quot;,&quot;nodeId&quot;:&quot;d5fa1c03-b0b1-4e72-bc88-7cd0ddc4118a&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-yaml">quality_certification: certified
data_contract_status: certified
quality_owner: finance-engineering
monitoring_required: true
sla_tier: gold
retention-policy: 1 year</code></pre></div><p>It does not change every few minutes. It is useful for discovery, governance, certification, access policies, platform automation, and downstream consumption.</p><p>The second category is operational quality state, something like:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;yaml&quot;,&quot;nodeId&quot;:&quot;0a81cbd6-ac72-44a9-addc-4f6e0524e4d3&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-yaml">last_freshness_check: failed
null_rate_customer_id: 0.08
row_count_anomaly_score: 0.91
failed_checks_count: 3
incident_id: MC-12345
last_run_id: dq_run_20260527_103000</code></pre></div><p>This information is also important.</p><p>But it is operational. It changes every time a check runs. It belongs to the monitoring layer, not necessarily to the table definition.</p><p>The industry often gets into trouble when it treats these two categories as the same thing.</p><p>They are not the same thing.</p><p></p><h1>What Data Quality support Delta Lake and Iceberg Already Provide</h1><p>Before inventing a new metadata model, it is worth looking at what open table formats already support.</p><p>Delta Lake and Apache Iceberg are not data quality platforms, but they do contain several building blocks that are useful for data quality.</p><h2>Delta Lake</h2><p>Delta Lake has a few native capabilities that are directly or indirectly related to data quality.</p><p>The most obvious one is constraint enforcement.</p><h3>Hard rules: constraint enforcement.</h3><p>Delta supports <em>NOT NULL</em> constraints. This is the simplest form of quality enforcement: a required field cannot be missing.</p><p>Delta also supports <em>CHECK</em> constraints, which allow teams to define rules such as:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;sql&quot;,&quot;nodeId&quot;:&quot;a5cfa69e-fd25-405b-8644-b9ab088bbad0&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-sql">ALTER TABLE finance.orders ADD CONSTRAINT valid_revenue
CHECK (revenue &gt;= 0);</code></pre></div><p>This is useful because the rule is enforced at write time. If bad data is written, the write fails.</p><p>That makes constraints stronger than a dashboard, stronger than a catalog tag, and stronger than a downstream alert.</p><h3>Soft expectations: key constraints.</h3><p>In some environments, especially when using Unity Catalog, teams can also define informational primary key and foreign key constraints. These are not always enforced in the same way as traditional relational database constraints, but they are still valuable metadata. They describe expected uniqueness and relationships between datasets.</p><h3>Technical signals: <code>data-skipping statistics.</code></h3><p>Delta also collects file-level statistics for data skipping. These statistics can include values such as minimum and maximum column values, and they help query engines avoid reading unnecessary files. These statistics are not designed as data quality indicators, but they can support quality-adjacent use cases.m For example, they can help answer questions like:</p><pre><code><code>Does this file contain unexpected value ranges?
Are some partitions empty?
Did a column suddenly stop appearing in newly written data?
Is the table layout still useful for common access patterns?</code></code></pre><p>But this is important: Delta statistics are primarily an optimization feature.</p><p>They are not a semantic data quality model.</p><p>So Delta gives us three useful layers:</p><pre><code><code>1. Hard rules        &#8594; NOT NULL, CHECK constraints
2. Soft expectations &#8594; informational keys, schema expectations
3. Technical signals &#8594; data-skipping statistics</code></code></pre><p>That is useful, but it is not the same as a full data quality platform.</p><p>Delta does not natively answer questions like:</p><pre><code><code>Is this table certified?
Did the freshness check fail this morning?
Is there an open incident?
Was the latest anomaly acknowledged?
Which team owns the failed check?
What was the quality score over the last 30 days?</code></code></pre><p>Those questions belong to the observability and governance layer.</p><p></p><h2>Apache Iceberg</h2><p>Apache Iceberg has a different but equally interesting metadata model.</p><p>Iceberg tracks table state through metadata files, snapshots, manifest lists, and manifest files. Each snapshot represents the state of the table at a point in time. This makes Iceberg strong for table evolution, reproducibility, rollback, and time travel.</p><h3>File-level metrics.</h3><p>Iceberg manifests track data files and include file-level metrics. Depending on the writer and table configuration, these metrics may include information such as:</p><pre><code><code>record counts
null value counts
lower and upper bounds
column sizes
value counts</code></code></pre><p>This is very useful metadata.</p><p>For data engineers, these metrics can help identify quality symptoms.</p><p>For example:</p><pre><code><code>A null count suddenly increases.
A partition has far fewer records than usual.
A timestamp upper bound is older than expected.
A numeric column has values outside the expected range.</code></code></pre><p>Again, though, Iceberg does not treat these as business-level data quality indicators. They are technical metadata used mostly for planning, pruning, and efficient reads.</p><h3>Rich table metadata.</h3><p>Iceberg(as well, as Delta lake) also supports custom table properties. This gives teams a simple place to attach stable metadata such as:</p><pre><code><code>quality_certification: certified
data_contract_status: certified
monitoring_required: true
quality_owner: finance-platform</code></code></pre><p>For richer metadata, Iceberg has Puffin files. Puffin is designed to store additional statistics or index-like metadata that does not fit naturally into Iceberg manifests.</p><p>This could theoretically support more advanced quality-related artifacts, especially if the industry wanted to standardize richer table-level statistics.</p><p>But even with Puffin, I would be careful.</p><p>IMHO, Puffin is a place for statistics and technical metadata. It should not become a dumping ground for every data quality run result, incident, alert, and failed check payload.</p><p></p><h2>The Important Distinction: stable quality metadata vs operational data quality indicators</h2><p>Delta and Iceberg already provide useful quality-adjacent primitives:</p><pre><code><code>Schema enforcement
Constraints
Snapshots
Table properties
Column statistics
File metrics
Metadata tables
Additional statistics artifacts</code></code></pre><p>These are valuable.</p><p>But they mostly answer structural and technical questions:</p><pre><code><code>Is this write valid?
What files belong to this table version?
What are the column-level file statistics?
What changed between snapshots?
What properties describe this table?</code></code></pre><p>They do not answer the higher-level trust questions users care about:</p><pre><code><code>Is this asset certified?
Is it safe for executive reporting?
Did the latest freshness check pass?
Is there an open data incident?
Is this table covered by a data contract?
Who is responsible for quality?
</code></code></pre><p><s>Delta and Iceberg need to become data quality systems.</s></p><p>I would say:</p><blockquote><p>Delta and Iceberg should expose stable quality metadata.</p></blockquote><p></p><h3>Operational Data Quality Results Do Not Belong in Table Metadata</h3><p>The obvious implementation is tempting.</p><p>Every time Monte Carlo, Soda, Great Expectations, dbt tests,  or a custom data quality framework runs, write the latest status back to the table metadata.</p><p>Something like:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;yaml&quot;,&quot;nodeId&quot;:&quot;304f2287-db86-4f50-bfb7-b97b538882cd&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-yaml">quality_status: failed
quality_score: 0.82
last_checked_at: 2026-05-27T10:30:00Z
failed_checks_count: 3
results_uri: s3://dq-results/finance_metrics/run_123.json</code></pre></div><p>For a small number of tables, this looks feasible. At platform scale, it becomes problematic. Data quality indicators are operational and time-sensitive. They change on every check:</p><p>Freshness can fail at 10:00 and recover at 10:15;</p><p>A volume anomaly can be detected in one run and disappear in the next;</p><p>A null-rate check can fail because of a temporary upstream delay;</p><p>Incidents can be opened, acknowledged, suppressed, escalated, or resolved;</p><p>If all of that is written directly into table metadata, the metadata layer becomes a high-churn operational store.</p><p>That creates several problems.</p><p>First, metadata history becomes noisy. Instead of capturing meaningful asset changes &#8212; schema updates, ownership changes, lifecycle transitions, contract certification &#8212; the table history gets filled with operational status updates.</p><p>Second, catalogs and sync systems are not designed to be incident event stores. They are optimized for discovery, governance, lineage, and relatively stable asset metadata. Constantly mutating properties across thousands of assets creates unnecessary load.</p><p>Third, consumers may read stale quality state. The data quality system may have the latest result, but the catalog sync may lag. The table property may show yesterday&#8217;s status. The BI tool may cache an older version. Now we have multiple versions of &#8220;truth.&#8221;</p><p>Fourth, large per-check payloads do not fit well into table properties. Detailed DQ output includes check names, thresholds, observed values, sample failures, incident links, owners, routing rules, and historical context.</p><p>Trying to squeeze this into table metadata turns metadata into an awkward JSON dump. That is not asset metadata anymore.</p><p>That is an operational log pretending to be metadata.</p><p>A data quality platform( Montecarlo, Soda, DBT tests, your custom data quality platform) should be the source of truth for operational quality state because it is designed to store check history, detect anomalies, route alerts, manage incidents, and provide debugging context. Catalogs and table metadata should only expose stable trust signals, such as certification or contract status, while detailed check results remain in the DQ platform.</p><h2>The Better Pattern: Stable Metadata + Operational Evidence</h2><p>The compromise I like is simple. </p><p>Put stable state in the asset metadata.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;yaml&quot;,&quot;nodeId&quot;:&quot;fd7c4228-8d3d-439a-a65d-432b3e5c138c&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-yaml">quality_certification: certified
data_contract_status: certified
quality_owner: data-platform
monitoring_required: true
sla_tier: gold</code></pre></div><p>Put run-level results in a dedicated system.</p><p>That could be Monte Carlo. It could be Soda Cloud. It could be Great Expectations Cloud. It could be a custom observability service. It could also be a sidecar Delta or Iceberg table if you want a queryable internal history, such as </p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;plaintext&quot;,&quot;nodeId&quot;:&quot;1bd2ee68-a50b-410f-9faf-e8a22473adcf&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-plaintext">dq_run_id
asset_id
checked_at
status
score
failed_checks_count
failed_checks
incident_id
results_uri
producer</code></pre></div><p>This gives us both things we need:</p><p>The asset remains clean and discoverable.</p><p>+</p><p>The operational history remains detailed and queryable.</p>]]></content:encoded></item><item><title><![CDATA[The Deletion Delusion: Your Modern Data Platform is Probably failing compliance]]></title><description><![CDATA[Modern data architectures vs GDPR]]></description><link>https://blog.dativo.io/p/the-deletion-delusion-your-modern</link><guid isPermaLink="false">https://blog.dativo.io/p/the-deletion-delusion-your-modern</guid><dc:creator><![CDATA[Sergey]]></dc:creator><pubDate>Mon, 20 Apr 2026 13:03:18 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!CPoY!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8f2ce7b6-bc58-41ad-b2a7-01114e05907e_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>There is a fundamental friction at the heart of modern data management: the widening chasm between legal fantasy and engineering reality. While your compliance/legal department assumes a "Right to Erasure" request is a simple SQL execution, every engineering lead knows the truth. Modern data platforms&#8212;built on &#8216;Big Data&#8217; and &#8216;Lakehouse&#8217; architectures&#8212;are optimized for append-heavy, read-intensive workloads. They were never designed for the selective, row-level mutations required by GDPR.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!CPoY!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8f2ce7b6-bc58-41ad-b2a7-01114e05907e_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!CPoY!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8f2ce7b6-bc58-41ad-b2a7-01114e05907e_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!CPoY!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8f2ce7b6-bc58-41ad-b2a7-01114e05907e_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!CPoY!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8f2ce7b6-bc58-41ad-b2a7-01114e05907e_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!CPoY!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8f2ce7b6-bc58-41ad-b2a7-01114e05907e_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!CPoY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8f2ce7b6-bc58-41ad-b2a7-01114e05907e_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8f2ce7b6-bc58-41ad-b2a7-01114e05907e_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2776797,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://blog.dativo.io/i/194278852?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8f2ce7b6-bc58-41ad-b2a7-01114e05907e_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!CPoY!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8f2ce7b6-bc58-41ad-b2a7-01114e05907e_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!CPoY!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8f2ce7b6-bc58-41ad-b2a7-01114e05907e_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!CPoY!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8f2ce7b6-bc58-41ad-b2a7-01114e05907e_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!CPoY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8f2ce7b6-bc58-41ad-b2a7-01114e05907e_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>While most/every company claims compliance, the underlying architecture of modern data lakes often makes true erasure a technical impossibility or hard(read expensive). Most organizations are operating under a "deletion delusion," where data is merely hidden from the application layer while remaining physically immutable in the depths of S3 or other storage.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://blog.dativo.io/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://blog.dativo.io/subscribe?"><span>Subscribe now</span></a></p><p></p><h2>The Financial Language of Privacy&#8212;ALE</h2><p>The biggest hurdle for privacy/platform engineering is securing a budget for a &#8220;Compliance Program.&#8221; To get the management&#8217;s attention, you must translate regulatory risk( practically - a city legend for management, until the are charged for lack of compliance) into <strong>Annual Loss Expectancy (ALE)</strong>. This isn&#8217;t just a compliance metric; it&#8217;s your <strong>ROSI (Return on Security Investment)</strong>.</p><p>The formula is: <strong>ALE = Probability of Failure (ARO) &#215; Impact (SLE)</strong>.</p><p>When calculating the <strong>Single Loss Expectancy (SLE)</strong>, don&#8217;t just look at the fine. You must include engineering remediation costs, legal fees, and the &#8220;compute surge&#8221; required to fix the data post-incident.</p><p><strong>The Financial Reality:</strong> If there is a <strong>2% chance</strong> of a material GDPR failure and the SLE (fine + remediation + legal) is <strong>&#8364;5,000,000</strong>, your ALE is:</p><p><strong>2% &#215; &#8364;5M = &#8364;100,000 / year</strong></p><p>If the cost to build a reliable automated shredder is &#8364;80,000, the program pays for itself. Without this quantitative model, you&#8217;re just an engineer asking for more &#8220;unproductive&#8221; budget.</p><h2>GDPR Assumes You Know What PII Is (You Don&#8217;t)</h2><p>The first point of failure isn&#8217;t a lack of intent; it&#8217;s a failure of <strong>Data Protection by Design</strong>. GDPR assumes you have a clear, static map of Personally Identifiable Information (PII). In a modern data platform, this is a <em>hallucination</em>.</p><p>Data classification is almost always incomplete, and in the complex web of modern pipelines, PII is transformed and re-created across 10+ systems. Even if you scrub a <code>user_id</code>, the user&#8217;s ghost persists via <strong>derived signals</strong> and <strong>indirect identifiers</strong>&#8212;session IDs, device fingerprints, and behavioral embeddings. AI systems can now re-identify individuals from data you never intentionally labeled as sensitive. If you can&#8217;t verify the location of every derived identifier, you are failing the accountability requirements.</p><div class="callout-block" data-callout="true"><p>Most companies are not non-compliant because they ignore GDPR. They are non-compliant because their architecture makes it (almost) impossible.</p></div><h2></h2><h2>"DELETE" is a Lie in the World of Immutable Files</h2><p>In traditional Relational Database  systems, deletion is a predictable row-level transaction. In the OLAP/Lakehouse world, your SQL <code>DELETE</code> is a lie. Because these systems rely on immutable columnar formats like Parquet, a delete command doesn&#8217;t remove data; it triggers a cascading failure of storage efficiency.</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!v6tZ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72abc99f-57ee-4cd0-adae-010692a0a2ae_1600x210.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!v6tZ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72abc99f-57ee-4cd0-adae-010692a0a2ae_1600x210.png 424w, https://substackcdn.com/image/fetch/$s_!v6tZ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72abc99f-57ee-4cd0-adae-010692a0a2ae_1600x210.png 848w, https://substackcdn.com/image/fetch/$s_!v6tZ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72abc99f-57ee-4cd0-adae-010692a0a2ae_1600x210.png 1272w, https://substackcdn.com/image/fetch/$s_!v6tZ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72abc99f-57ee-4cd0-adae-010692a0a2ae_1600x210.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!v6tZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72abc99f-57ee-4cd0-adae-010692a0a2ae_1600x210.png" width="1456" height="191" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/72abc99f-57ee-4cd0-adae-010692a0a2ae_1600x210.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:191,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:83669,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://blog.dativo.io/i/194278852?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72abc99f-57ee-4cd0-adae-010692a0a2ae_1600x210.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!v6tZ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72abc99f-57ee-4cd0-adae-010692a0a2ae_1600x210.png 424w, https://substackcdn.com/image/fetch/$s_!v6tZ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72abc99f-57ee-4cd0-adae-010692a0a2ae_1600x210.png 848w, https://substackcdn.com/image/fetch/$s_!v6tZ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72abc99f-57ee-4cd0-adae-010692a0a2ae_1600x210.png 1272w, https://substackcdn.com/image/fetch/$s_!v6tZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72abc99f-57ee-4cd0-adae-010692a0a2ae_1600x210.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p>Behind a &#8220;Logical Table: Deleted&#8221; checkmark, a standard delete command actually causes:</p><ul><li><p><strong>Massive I/O &amp; Compute Spikes:</strong> The system must scan, filter, rewrite, and replace entire file groups.</p></li><li><p><strong>S3 Tier Promotion:</strong> Deletion jobs often drag data from &#8220;Cold&#8221; storage tiers back to &#8220;Hot,&#8221; spiking your monthly cloud bill while failing to actually purge the data.</p></li><li><p><strong>Passive Propagation:</strong> Deleted data persists in:</p><ul><li><p>&#10060; <strong>Delta/Iceberg History:</strong> Transaction logs maintain the state for &#8220;time travel.&#8221;</p></li><li><p>&#10060; <strong>S3/Cloud Backups:</strong> Data lives on in secondary storage and snapshots.</p></li><li><p>&#10060; <strong>Stale Aggregates:</strong> Downstream summaries still contain the mathematical influence of the deleted records.</p></li></ul></li></ul><p></p><p></p><p><strong>The AI Black Hole&#8212;Logs, Embeddings, and the Unknown</strong></p><p>The most dangerous tier of the deletion hierarchy is the &#8220;AI/Unknown&#8221; tier. Modern humanity are currently feeding massive amounts of PII into Large Language Models (LLMs), creating a trail of logs and high-dimensional embeddings that are mathematically impossible to &#8220;un-learn&#8221; or selectively purge.</p><p>Passive cleanup is no longer enough. We need an <strong><a href="https://github.com/dativo-io/talon">AI Control Plane</a></strong> that moves privacy to active runtime enforcement. This control plane must at least:</p><ul><li><p>Understand data sensitivity at the point of ingestion.</p></li><li><p>Enforce policies at the prompt/inference level.</p></li><li><p>Provide verifiable evidence of erasure across the entire AI lifecycle.</p><p></p></li></ul><p></p><p><strong>Conclusion: Beyond the Checkbox</strong></p><p>True compliance requires a shift from &#8220;passive compliance&#8221; to &#8220;active privacy engineering.&#8221; A defensible strategy involves a phased rollout based on the <strong> </strong>priority hierarchy, such as for example:</p><ol><li><p>Mandatory Right to Erasure (DDR) and Account Deletion.</p></li><li><p>Content deletion and retroactive backfills.</p></li><li><p>Addressing indirect identifiers (such as session IDs), full downstream propagation, and auditability.</p></li></ol><p><strong>Final Thought:</strong> Is your current &#8220;DELETE&#8221; button actually removing data, or is it just hiding it from view? The <strong><a href="https://www.edpb.europa.eu/coordinated-enforcement-framework_en">2026 European Data Protection Board&#8217;s report</a></strong> makes it clear: regulators are no longer accepting &#8220;technical difficulty&#8221; or &#8220;immutable architecture&#8221; as valid excuses for data persistence. They are specifically targeting gaps in backup erasure and failed anonymization. If your data platform treats deletion as an afterthought, you aren&#8217;t compliant&#8212;you&#8217;re just lucky. For now.</p>]]></content:encoded></item><item><title><![CDATA[If we redact PII before the model sees the prompt, can we still preserve enough context for good reasoning?]]></title><description><![CDATA[Same privacy boundary. Better answers. Measured across 200 A/B prompts.]]></description><link>https://blog.dativo.io/p/if-we-redact-pii-before-the-model</link><guid isPermaLink="false">https://blog.dativo.io/p/if-we-redact-pii-before-the-model</guid><dc:creator><![CDATA[Sergey]]></dc:creator><pubDate>Fri, 27 Mar 2026 21:41:22 GMT</pubDate><enclosure url="https://images.unsplash.com/photo-1764924671797-8d546240552b?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2NHx8Y29udGV4dHxlbnwwfHx8fDE3NzQ2NDY1ODZ8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://images.unsplash.com/photo-1764924671797-8d546240552b?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2NHx8Y29udGV4dHxlbnwwfHx8fDE3NzQ2NDY1ODZ8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://images.unsplash.com/photo-1764924671797-8d546240552b?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2NHx8Y29udGV4dHxlbnwwfHx8fDE3NzQ2NDY1ODZ8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1764924671797-8d546240552b?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2NHx8Y29udGV4dHxlbnwwfHx8fDE3NzQ2NDY1ODZ8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1764924671797-8d546240552b?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2NHx8Y29udGV4dHxlbnwwfHx8fDE3NzQ2NDY1ODZ8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1764924671797-8d546240552b?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2NHx8Y29udGV4dHxlbnwwfHx8fDE3NzQ2NDY1ODZ8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw"><img src="https://images.unsplash.com/photo-1764924671797-8d546240552b?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2NHx8Y29udGV4dHxlbnwwfHx8fDE3NzQ2NDY1ODZ8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" width="4000" height="6000" data-attrs="{&quot;src&quot;:&quot;https://images.unsplash.com/photo-1764924671797-8d546240552b?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2NHx8Y29udGV4dHxlbnwwfHx8fDE3NzQ2NDY1ODZ8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:6000,&quot;width&quot;:4000,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Wooden sculpture of a woman reading a book.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Wooden sculpture of a woman reading a book." title="Wooden sculpture of a woman reading a book." srcset="https://images.unsplash.com/photo-1764924671797-8d546240552b?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2NHx8Y29udGV4dHxlbnwwfHx8fDE3NzQ2NDY1ODZ8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1764924671797-8d546240552b?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2NHx8Y29udGV4dHxlbnwwfHx8fDE3NzQ2NDY1ODZ8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1764924671797-8d546240552b?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2NHx8Y29udGV4dHxlbnwwfHx8fDE3NzQ2NDY1ODZ8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1764924671797-8d546240552b?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2NHx8Y29udGV4dHxlbnwwfHx8fDE3NzQ2NDY1ODZ8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Photo by <a href="https://unsplash.com/@le_y0u">You Le</a> on <a href="https://unsplash.com">Unsplash</a></figcaption></figure></div><p></p><p>Most teams still talk about privacy and model quality as if you can only have one.</p><p>Either you protect sensitive data, or you preserve enough context for the model to be useful.</p><p>That tradeoff sounds intuitive. It is also too simplistic.</p><p>Ideally with tools like <a href="https://github.com/dativo-io/talon">Talon</a>, raw PII never reaches the model. But there is a big difference between removing PII and removing meaning.</p><p>A flat placeholder like <code>[PHONE]</code> protects privacy, but it also hides the one thing the model may actually need to answer correctly: is this a German number, a Polish number, or a French one?</p><p>So I tested a different approach.</p><p>Instead of sending raw personal data to the model, Talon can replace it with structured placeholders that preserve only safe, task-relevant semantics. Not the original value. Just the minimum useful context.</p><p>And when I compared that enriched approach against legacy type-only redaction, the enriched version won in both evaluation runs.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://blog.dativo.io/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://blog.dativo.io/subscribe?"><span>Subscribe now</span></a></p><p></p><h2>TL;DR</h2><p>I ran two 100-prompt A/B evaluations on Dativo Talon to test a simple question: if we redact PII before the model sees input, can we still preserve enough meaning for useful reasoning? Answer: <strong>yes</strong>. Enriched redaction (semantic placeholders) beat legacy type-only redaction in both runs, especially on attribute-dependent tasks like country routing and payment-method decisions.</p><h2>The setup</h2><p>I tested two variants.</p><p><strong>Variant A: legacy redaction</strong><br>The model sees flat placeholders such as <code>[PERSON]</code>, <code>[EMAIL]</code>, <code>[PHONE]</code>, <code>[IBAN]</code>.</p><p><strong>Variant B: enriched redaction</strong><br>The model sees structured placeholders such as:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;xml&quot;,&quot;nodeId&quot;:&quot;4dcd9da6-a431-45b7-b401-ebcf4cd7c036&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-xml">&lt;PII type=&#8221;phone&#8221; country_code=&#8221;PL&#8221;/&gt;</code></pre></div><p>or</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;xml&quot;,&quot;nodeId&quot;:&quot;51a67038-9818-4b39-b241-14e46c056955&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-xml">&lt;PII type=&#8221;email&#8221; domain_type=&#8221;free&#8221;/&gt;</code></pre></div><p>In both cases, raw PII is removed before model input.</p><p>This is not a comparison between &#8220;private&#8221; and &#8220;non-private.&#8221; Both variants keep raw personal data away from the model. The only difference is whether the model still receives safe semantic hints that matter for the task.</p><p>I ran two full experiments:</p><ul><li><p><strong>Run 1:</strong> <code>gpt-4o-mini</code>, 100 prompts.  <a href="https://gist.github.com/sergeyenin/3e90542d43c8c58d3bf12e0743ae10dd">See full logs</a></p><p><br></p></li><li><p><strong>Run 2:</strong> <code>gpt-4o</code>, 100 prompts. <a href="https://gist.github.com/sergeyenin/218ce51662482288e1d24b0defe77edb">See full logs</a><br></p></li></ul><p>Each prompt was scored on four dimensions, from 1 to 10:</p><ol><li><p>attribute reasoning(&#8220;context&#8221;)</p></li><li><p>utility preservation</p></li><li><p>semantic coherence</p></li><li><p>helpfulness<br></p></li></ol><p>So each prompt had a maximum total score of <strong>40</strong>.</p><p></p><h2>What happened</h2><h3>Run 1 &#8212; <code>gpt-4o-mini</code> (N=100)</h3><ul><li><p><strong>A mean total:</strong> 23.87</p></li><li><p><strong>B mean total:</strong> 28.58</p></li><li><p><strong>Mean delta:</strong> +4.71</p></li><li><p><strong>Numeric wins:</strong> B 59, A 19, ties 22</p></li></ul><h3>Run 2 &#8212; <code>gpt-4o</code> (N=100)</h3><ul><li><p><strong>A mean total:</strong> 23.5</p></li><li><p><strong>B mean total:</strong> 30.2</p></li><li><p><strong>Mean delta:</strong> +6.7</p></li><li><p><strong>Numeric wins:</strong> B 63, A 26, ties 11</p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!7OoY!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F35d11069-6839-4fff-9f62-dcc17ecdfa60_956x216.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!7OoY!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F35d11069-6839-4fff-9f62-dcc17ecdfa60_956x216.png 424w, https://substackcdn.com/image/fetch/$s_!7OoY!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F35d11069-6839-4fff-9f62-dcc17ecdfa60_956x216.png 848w, https://substackcdn.com/image/fetch/$s_!7OoY!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F35d11069-6839-4fff-9f62-dcc17ecdfa60_956x216.png 1272w, https://substackcdn.com/image/fetch/$s_!7OoY!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F35d11069-6839-4fff-9f62-dcc17ecdfa60_956x216.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!7OoY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F35d11069-6839-4fff-9f62-dcc17ecdfa60_956x216.png" width="956" height="216" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/35d11069-6839-4fff-9f62-dcc17ecdfa60_956x216.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:216,&quot;width&quot;:956,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:38688,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://blog.dativo.io/i/192355842?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F35d11069-6839-4fff-9f62-dcc17ecdfa60_956x216.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!7OoY!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F35d11069-6839-4fff-9f62-dcc17ecdfa60_956x216.png 424w, https://substackcdn.com/image/fetch/$s_!7OoY!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F35d11069-6839-4fff-9f62-dcc17ecdfa60_956x216.png 848w, https://substackcdn.com/image/fetch/$s_!7OoY!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F35d11069-6839-4fff-9f62-dcc17ecdfa60_956x216.png 1272w, https://substackcdn.com/image/fetch/$s_!7OoY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F35d11069-6839-4fff-9f62-dcc17ecdfa60_956x216.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p>Enriched redaction wins in both runs.<br>The strongest lift is exactly where it should be: <strong>attribute reasoning</strong>.</p><h2>Where enrichment helped most</h2><p>Not all semantic attributes are equally valuable.</p><p>The strongest gains came from attributes that directly affect routing, jurisdiction, or payment behavior.</p><h3>Strongest lift</h3><ul><li><p><strong>PHONE &#8594; </strong><code>country_code</code></p></li><li><p><strong>IBAN &#8594; </strong><code>country_code</code></p></li></ul><p>This was the clearest signal in both runs.</p><p>If the model sees only <code>[PHONE]</code>, it cannot reliably decide whether the request belongs with Germany, Poland, France, or another support flow.</p><p>If it sees <code>&lt;PII type="phone" country_code="DE"/&gt;</code>, that ambiguity disappears without exposing the original number.</p><p>The same pattern showed up for IBANs. A country code is often enough to reason about SEPA, local handling, or country-specific banking logic.</p><h3>Moderate lift</h3><ul><li><p><strong>PERSON &#8594; </strong><code>gender</code></p></li><li><p><strong>EMAIL &#8594; </strong><code>domain_type</code></p></li></ul><p>These still helped, but less dramatically.</p><p>That also makes sense.</p><p>A corporate email domain or a gendered title can improve the answer, but the downstream task is often less deterministic than country-based routing. The model can sometimes get close with generic language even without the attribute.</p><h3>Weakest area</h3><ul><li><p><strong>LOCATION &#8594; </strong><code>scope</code> such as city, region, or country</p></li></ul><p>This was the least convincing category.</p><p>The issue was not necessarily the enrichment itself. It was the prompt design.</p><p>Too many location prompts could still be answered with generic legal boilerplate. If a question does not force the model to actually use the distinction between city, region, and country, then that attribute will not show its value.</p><p>So this is less &#8220;scope does not help&#8221; and more &#8220;the benchmark did not pressure-test scope hard enough.&#8221;</p><h2>What this means in practice</h2><p>This is the production lesson.</p><p>Most teams fall into one of two bad patterns.</p><p>The first is to send raw prompts with PII to the model and hope governance happens somewhere later.</p><p>The second is to over-redact everything into useless placeholder soup and then act surprised when the model starts guessing.</p><p>Neither is a good long-term design.</p><p>The better path is narrower and more disciplined:</p><ul><li><p>remove raw PII before the model sees it</p></li><li><p>preserve only the minimum safe semantics needed for reasoning</p></li><li><p>decide those semantics through policy</p></li><li><p>record evidence of what the model actually saw</p></li></ul><p>That is the operating model Talon is built around, and this evaluation supports it.</p><h2>Edge cases worth being honest about</h2><p>The result is strong, but not perfect.</p><p>A few caveats matter.</p><p><strong>1. Prompt quality still varied</strong></p><p>Some prompts were genuinely attribute-dependent. Others were only loosely so.</p><p>That matters because if a prompt can be answered with generic common sense, the benchmark becomes less discriminative.</p><p><strong>2. Judge behavior still has style bias</strong></p><p>In a few cases, longer and more generic answers scored surprisingly well, even when a shorter answer was more precise.</p><p>That is a familiar problem in LLM-as-judge evaluations.</p><p><strong>3. Order effects were more visible on the smaller model</strong></p><p>I saw a bit more sensitivity in the <code>gpt-4o-mini</code> run than in the <code>gpt-4o</code> run.</p><p>Not enough to change the direction of the result, but enough to keep in mind.</p><p><strong>4. Location-scope prompts need to be redesigned</strong></p><p>This was the weakest benchmark segment and the one I would trust least in its current form.</p><p>So yes, the result is real. But some parts of the evaluation are stronger than others.</p><p></p><h2>What about cost?</h2><p>This is usually the first practical objection.</p><p>Semantic placeholders are longer. So do they make inference meaningfully more expensive?</p><p>In these runs, not in a way that mattered.</p><ul><li><p>In the <code>gpt-4o-mini</code> run, Variant B was slightly cheaper overall.</p></li><li><p>In the <code>gpt-4o</code> run, Variant B was materially cheaper in observed run cost.</p></li></ul><p>That does not mean enriched placeholders are inherently cheaper token-for-token. They are not.</p><p>It means end-to-end cost is dominated by full model behavior, especially output length and answer shape, not just placeholder size.</p><p>So the real takeaway is simpler:</p><p><strong>The quality gain was clear, and the added redaction structure did not create a practical cost penalty.</strong></p><p></p><h2>Why This Matters for Production</h2><p>Most teams still run one of two broken patterns:</p><ol><li><p>send raw prompts with PII and hope policy catches up later, or</p></li><li><p>over-redact into useless <code>[TYPE]</code> soup and lose utility.</p></li></ol><p>There is a better middle path:</p><ul><li><p>remove raw PII from model input</p></li><li><p>preserve a minimal, safe semantic layer</p></li><li><p>apply policy-as-code to which attributes are allowed</p></li><li><p>keep evidence for what the model actually saw</p></li></ul><p>That is what this experiment validates.</p><h2>What we are changing next in Talon</h2><p>Based on these runs, here is what I would do next:</p><ol><li><p><strong>Keep enriched redaction as the default</strong> for supported models.</p></li><li><p><strong>Improve the location-scope benchmark</strong>, especially prompt quality and dictionaries.</p></li><li><p><strong>Use a separate judge model</strong> and add a small human-reviewed subset.</p></li><li><p><strong>Run multi-seed evaluations with confidence intervals</strong> in CI.</p></li></ol><p>The core result is already useful. But the next version of the benchmark should be harder, cleaner, and more defensible.</p><h2>Final Thought</h2><p>Privacy-preserving AI does not need to be blind AI.</p><p>If your redaction layer removes everything useful, the model will guess.<br>If your redaction layer keeps safe semantics, the model can still reason.</p><p>This is not a theoretical point. We measured it across 200 prompt pairs.</p><p>No raw PII to the model. Better answers anyway.</p>]]></content:encoded></item><item><title><![CDATA[Why AI Agents Work in Demos but Fail in Production]]></title><description><![CDATA[Unless you&#8217;re doing it right&#8212;and right now, almost nobody is]]></description><link>https://blog.dativo.io/p/why-ai-agents-work-in-demos-but-fail</link><guid isPermaLink="false">https://blog.dativo.io/p/why-ai-agents-work-in-demos-but-fail</guid><dc:creator><![CDATA[Sergey]]></dc:creator><pubDate>Wed, 18 Mar 2026 10:35:50 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!fYSI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2dcf8ed5-0992-4f6a-9aa9-2755bb4bee88_1024x1536.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!fYSI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2dcf8ed5-0992-4f6a-9aa9-2755bb4bee88_1024x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!fYSI!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2dcf8ed5-0992-4f6a-9aa9-2755bb4bee88_1024x1536.png 424w, https://substackcdn.com/image/fetch/$s_!fYSI!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2dcf8ed5-0992-4f6a-9aa9-2755bb4bee88_1024x1536.png 848w, https://substackcdn.com/image/fetch/$s_!fYSI!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2dcf8ed5-0992-4f6a-9aa9-2755bb4bee88_1024x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!fYSI!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2dcf8ed5-0992-4f6a-9aa9-2755bb4bee88_1024x1536.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!fYSI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2dcf8ed5-0992-4f6a-9aa9-2755bb4bee88_1024x1536.png" width="1024" height="1536" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2dcf8ed5-0992-4f6a-9aa9-2755bb4bee88_1024x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1536,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2457871,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://blog.dativo.io/i/191348674?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2dcf8ed5-0992-4f6a-9aa9-2755bb4bee88_1024x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!fYSI!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2dcf8ed5-0992-4f6a-9aa9-2755bb4bee88_1024x1536.png 424w, https://substackcdn.com/image/fetch/$s_!fYSI!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2dcf8ed5-0992-4f6a-9aa9-2755bb4bee88_1024x1536.png 848w, https://substackcdn.com/image/fetch/$s_!fYSI!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2dcf8ed5-0992-4f6a-9aa9-2755bb4bee88_1024x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!fYSI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2dcf8ed5-0992-4f6a-9aa9-2755bb4bee88_1024x1536.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">The AI silver bullet don&#8217;t exist as well :(</figcaption></figure></div><p>The current obsession with model &#8220;intelligence&#8221; is a failure of engineering discipline. CTOs and Senior Architects are chasing leaderboard scores like teenagers chasing fashion trends, only to watch their multi-agent systems (MAS) implode the moment they hit production.</p><p>The primary cause of multi-agent failure isn&#8217;t &#8220;weak models&#8221; - every week, a new model tops another benchmark. Every month, another company claims its agents are more autonomous, more intelligent, more capable. And yet the same thing keeps happening in production: Multi-agent systems look impressive in demos, then fall apart under real load. Not because the models are weak, but because the framework is.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://blog.dativo.io/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption"></p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>The industry is solving the wrong problem. We are building systems as if LLMs are deterministic components, ignoring the reality that uncertainty at every agent handoff creates a multiplicative decay in reliability. High-performing models in a controlled notebook are a demo-day vanity metric. In the real world, &#8220;stochastic hope&#8221; is an engineering liability. Reliable agentic systems are not built by choosing better models; they are built by engineering rigid data boundaries and treating the entire system as a distributed data pipeline.</p><p>A multi-agent system is not &#8220;a group of smart agents working together&#8221;. It is a <strong>distributed pipeline of untrusted intermediate states</strong>.</p><p></p><h1>Software Engineering is dead, longs live the software engineering</h1><h2><strong>Agentic Systems as Probabilistic Pipelines</strong></h2><p>When you wire multiple agents together, you are building a series-system pipeline <a href="https://www.oreilly.com/radar/the-hidden-cost-of-agentic-failure/?utm_medium=email&amp;utm_source=platform+b2c&amp;utm_campaign=rediscover&amp;utm_content=canceled+20260317">governed by </a><strong><a href="https://www.oreilly.com/radar/the-hidden-cost-of-agentic-failure/?utm_medium=email&amp;utm_source=platform+b2c&amp;utm_campaign=rediscover&amp;utm_content=canceled+20260317">Lusser&#8217;s Law</a> </strong>.  In reliability engineering, this is the logic of a <strong>series system</strong>: if a workflow requires multiple sequential steps, the total success rate is the product of the success rates of the individual steps. </p><p>A single agent can look excellent in isolation&#8212;and that is exactly the trap. A model with 98% task accuracy sounds production-ready, but production systems are judged end-to-end.</p><p>In reliability engineering, a sequential workflow&#8217;s success is the product of the reliability of each step. If each hop succeeds with probability $p$, then an $n$-step workflow succeeds with probability:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;P(\\text{system success}) = p^n&quot;,&quot;id&quot;:&quot;HFEVPHFUJV&quot;}" data-component-name="LatexBlockToDOM"></div><p>Even with a stronger agents backed by better models, the decay is sharp:</p><ul><li><p>1 agent at 98%     &#8594;   Total success: 98.0%</p></li><li><p>5 agents at 98%   &#8594;    Total success:  90.4%</p></li><li><p>10 agents at 98% &#8594;    Total success: 81.7%</p></li></ul><p>One bad output does not just fail locally; it becomes state. The next agent reads it, trusts it, and reasons on top of it. A hallucinated tool response doesn&#8217;t just reduce the chance of success at one step; it <strong>poisons the steps that follow</strong>.</p><p>On the production floor, these mathematical decays manifest as destructive pathologies that eat your budget and kill your uptime:</p><ul><li><p><strong>Silent Schema Drift:</strong> A model outputs a slightly malformed JSON or an unexpected type. Without a validation gate, this corrupted state propagates downstream. Subsequent agents then condition their &#8220;reasoning&#8221; on garbage, leading to catastrophic cascades.</p></li><li><p><strong>Hallucinated Tool Outputs:</strong> Agents often condition their next action on a false return from an unvalidated tool call. Without a control plane to verify the return, the error remains invisible until the system produces a hallucinated final result.</p></li><li><p><strong>Unvalidated Handoffs:</strong> This is the peak of &#8220;stochastic hope&#8221;&#8212;passing raw strings between agents and praying the next model correctly parses the intent. It is the architectural equivalent of using <code>eval()</code> on untrusted user input.</p></li><li><p><strong>Operational Death Spirals:</strong> Recursive reasoning loops where supervisors fail to reach a terminal state. These loops consume thousands of tokens in seconds, draining API budgets without making an inch of progress toward the objective.</p></li></ul><p>This is the key mindset shift: a multi-agent system is not &#8220;a set of smart models collaborating.&#8221; It is a <strong>distributed pipeline of untrusted intermediate states</strong>.</p><p>Once you see it that way, the engineering answer becomes obvious. You do not solve the problem by making every model slightly better. You solve it by inserting <strong>contracts, validation gates, and control points</strong> so the system can survive when one hop is wrong.</p><h3><strong>The Analogy: MAS are Untyped Distributed Pipelines</strong></h3><p>The modern multi-agent stack looks a lot like the messy early days of data engineering. We are passing intermediate state between stages without schemas, relying on downstream logic to &#8220;figure out&#8221; malformed upstream outputs.</p><p>This isn&#8217;t an AI problem&#8212;it&#8217;s a <strong>data reliability problem</strong>. Until we treat agentic handoffs as formal contracts, these systems will never scale.</p><h1><strong>The Missing Layer: Contracts , Validation and Control</strong></h1><p>The only way to break the multiplicative decay of Lusser&#8217;s Law is to introduce gates. By verifying an output before it reaches the next agent, you change the &#8220;reliability math.&#8221;</p><p>The <strong>Effective Probability formula</strong> </p><p></p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;p_{\\text{effective}} = p + (1 - p) \\cdot v&quot;,&quot;id&quot;:&quot;ZOKOFDPYDM&quot;}" data-component-name="LatexBlockToDOM"></div><p>shows us the way out. By applying a validation catch rate (<em>v</em>), you recover failures before they propagate. A 98% accurate agent with a 90% validation catch rate becomes 99.8% effective. Over 10 hops, that&#8217;s the difference between an 81.7% failure-prone system and a 98% stable one.</p><p></p><h1>A recipe for good life with AI Agents</h1><h2>Recipe 1: Pydantic + Instructor for Handoff Contracts</h2><p>The first job is to stop bad state from propagating. That means validation gates.</p><p><strong>Rule: Never pass raw LLM output to the next agent</strong>. Use <a href="https://docs.pydantic.dev/latest/">Pydantic</a> to define a contract and Instructor to force the model to satisfy it.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;373b9313-24ae-4023-bccb-aaf2f587c471&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">import instructor
from openai import OpenAI
from pydantic import BaseModel, Field, model_validator
from typing import Literal, Optional

client = instructor.from_openai(OpenAI(), max_retries=2)

class TicketDecision(BaseModel):
    action: Literal["approve", "reject", "escalate"]
    ticket_id: str
    risk_score: float = Field(ge=0, le=1)
    reason: str = Field(min_length=20, max_length=500)
    approver_id: Optional[str] = None

    @model_validator(mode="after")
    def enforce_business_rules(self):
        if self.action == "approve" and self.risk_score &gt; 0.7:
            raise ValueError("high-risk tickets cannot be auto-approved")
        return self</code></pre></div><p>This shifts the system from &#8220;trust the prompt&#8221; to &#8220;trust only validated state&#8221;.</p><h2>Recipe 2: Best-of-N + Controlled Ranking</h2><p>Validation prevents malformed outputs, but it doesn&#8217;t tell you which valid output is <em>best</em>. For complex tasks, use <strong>Best-of-N generation</strong> followed by a ranking step (like <a href="https://openpipe.ai/blog/ruler">RULER</a>).</p><ol><li><p><strong>Generate:</strong> Create 4 candidates.</p></li><li><p><strong>Validate:</strong> Filter out any that fail the Pydantic schema.</p></li><li><p><strong>Rank:</strong> Use a judge model to pick the winner based on a specific rubric (correctness, policy, clarity).</p></li></ol><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;d2251207-6ded-4468-a66a-eccdb187a28e&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from typing import List
import instructor
from pydantic import BaseModel

# 1. Define the Contract
class AnalysisReport(BaseModel):
    summary: str
    sentiment: str
    confidence_score: float

# 2. Generate N Candidates
def generate_candidates(task: str, n: int = 4) -&gt; List[AnalysisReport]:
    candidates = []
    for _ in range(n):
        try:
            # Each call is a stochastic draw
            res = client.chat.completions.create(
                model="gpt-4o",
                response_model=AnalysisReport,
                messages=[{"role": "user", "content": task}]
            )
            candidates.append(res)
        except Exception:
            continue # Skip malformed candidates
    return candidates

# 3. Apply RULER (The Judge)
def ruler_rank(candidates: List[AnalysisReport], task: str) -&gt; AnalysisReport:
    # We ask a stronger model to act as the 'RULER' judge
    judge_prompt = f"""
    Task: {task}
    Candidates: {candidates}
    
    Rank these candidates based on:
    1. Depth of insight in the 'summary'
    2. Alignment between 'sentiment' and 'summary'
    3. Realistic 'confidence_score' (avoid overconfidence)
    
    Return only the best candidate's index.
    """
    # Logic to select the winner based on the judge's decision
    # ...
    return candidates[winner_index]</code></pre></div><p></p><h2>Recipe 3: Talon for Bounded Search and Budget Control</h2><p>Local validation is a start, but production reliability requires an <strong>external control gateway</strong>. This is where <strong><a href="https://github.com/dativo-io/talon">Talon</a></strong> fits.</p><p>Search-based reliability (like Best-of-N or multi-step reasoning) is a double-edged sword. Without a control plane, a "smart" supervisor might trigger an infinite loop of retries, re-rankings, and judge calls. This leads to "test-time bankruptcy," where a single user request consumes hundreds of dollars in tokens. Talon places hard caps on the reasoning process itself.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;yaml&quot;,&quot;nodeId&quot;:&quot;b152a267-ab15-402b-937c-c21ac86f68fd&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-yaml">policies:
  session_limits:
    max_cost: 2.50          # Absolute dollar cap per trace
    max_candidates: 4       # Limit Best-of-N generation width
    max_judge_calls: 2      # Limit the number of re-evaluations</code></pre></div><h2>Recipe 4: Talon as the Validated Commit Boundary</h2><p>In production, you should never give an LLM &#8220;raw&#8221; write access to your database or APIs. An agent is a stochastic engine; if it hallucinates a <code>DELETE</code> flag or an unbounded <code>LIMIT</code>, the damage is instantaneous. Talon acts as a &#8220;Commit Wrapper,&#8221; forcing every tool call to pass through a deterministic governance layer before execution.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;yaml&quot;,&quot;nodeId&quot;:&quot;61276251-1b99-4295-a5bb-e4dac109417c&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-yaml">tool_policies:
  update_customer_records:
    max_row_count: 100            # Block "Update All" hallucinations
    require_dry_run: true         # Force a simulation first
    forbidden_argument_values:
      mode: ["truncate", "drop"]  # Block destructive operations
    arguments:
      query: redact               # Strip PII from logs/traces
    timeout: "15s"                # Kill runaway tool executions</code></pre></div><p></p><h2>Recipe 5: Talon Idempotency to Stop Duplicate Side Effects</h2><p>Retries are a requirement for reliability, but they are dangerous for side-effecting tools like sending emails or charging credit cards. If an upstream planner fails and retries the entire sequence, you risk executing the same action twice. Talon tracks the intent of a tool call, ensuring that repeated calls with the same parameters do not trigger duplicate external actions.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;yaml&quot;,&quot;nodeId&quot;:&quot;2bcfde40-9bc7-49f2-ae2c-2587ed62b1c6&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-yaml">tool_governance:
  send_notification_email:
    idempotency_key: "request_id" # Link to the unique session ID
    cache_ttl: "24h"              # Prevent double-send within a window
    on_duplicate: "return_cached" # Return the original success response
    strict_mode: true             # Fail if idempotency cannot be verified</code></pre></div><p></p><h1><strong>Engineering by Design</strong></h1><p>Engineering reliable massive agentic systems requires two fundamental shifts in leadership perspective:</p><ol><li><p><strong>From &#8220;Models are smart&#8221; to &#8220;Systems must be safe under uncertainty.&#8221;</strong> Assume the model will fail. Build the safety net first.</p></li><li><p><strong>From &#8220;Prompt Engineering&#8221; to &#8220;System Architecture.&#8221;</strong> Reliability is a function of boundary enforcement and budget control, not the phrasing of a system prompt.</p></li></ol><p>Agentic systems do not become reliable by chance; they become reliable by design. Scalable AI is a data engineering challenge involving state management, contract enforcement, and cost governance. By implementing strict validation boundaries, using Talon as a control plane, and amortizing intelligence through Reinforcement Learning, you move from &#8220;stochastic hope&#8221; to production-ready infrastructure. Reliability is a choice. Make it.</p><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://blog.dativo.io/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Data, Engineering, and Beyond! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[LLM APIs Have No Seatbelts. I Built One.]]></title><description><![CDATA[Why People Are Putting a Reverse Proxy in Front of Their AI Traffic]]></description><link>https://blog.dativo.io/p/llm-apis-have-no-seatbelts-i-built</link><guid isPermaLink="false">https://blog.dativo.io/p/llm-apis-have-no-seatbelts-i-built</guid><dc:creator><![CDATA[Sergey]]></dc:creator><pubDate>Sat, 28 Feb 2026 15:24:11 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Pehe!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9850867b-aa32-4812-9895-a097dfc96032_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>TL;DR</strong>: LLM APIs don&#8217;t ship with the controls you&#8217;d expect from any other piece of infrastructure &#8212; no per-caller auth, no tool restrictions, no cost ceiling. I got burned (meeting invitation I did not want to sent, $385 weekend loop, private information, like IBAN, in plaintext) and built a reverse proxy that fixes it. Here&#8217;s how three features work in practice, with the exact configs I run.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Pehe!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9850867b-aa32-4812-9895-a097dfc96032_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Pehe!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9850867b-aa32-4812-9895-a097dfc96032_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!Pehe!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9850867b-aa32-4812-9895-a097dfc96032_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!Pehe!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9850867b-aa32-4812-9895-a097dfc96032_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!Pehe!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9850867b-aa32-4812-9895-a097dfc96032_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Pehe!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9850867b-aa32-4812-9895-a097dfc96032_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9850867b-aa32-4812-9895-a097dfc96032_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:3795710,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://blog.dativo.io/i/189337071?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9850867b-aa32-4812-9895-a097dfc96032_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Pehe!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9850867b-aa32-4812-9895-a097dfc96032_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!Pehe!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9850867b-aa32-4812-9895-a097dfc96032_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!Pehe!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9850867b-aa32-4812-9895-a097dfc96032_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!Pehe!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9850867b-aa32-4812-9895-a097dfc96032_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p>It was a Sunday afternoon. I was building a personal scheduling agent &#8212; the kind that reads your calendar, finds gaps, and books meetings automatically. Super useful for coordinating squash or catching up with friends. I&#8217;d been hacking on it for a couple of days and wanted to test the full flow end-to-end. </p><p>I needed test contacts which would react, so I exported a few from my phone &#8212; figured I&#8217;d use people I actually know. My friends Marek and Tomek, and a couple of others. I told the agent to &#8220;book some test meetings for next week&#8221; and went to make coffee.</p><p>By the time I got back, it had sent real calendar invites. To all of them. For a meeting titled &#8220;Test Meeting 3&#8221; with no agenda, no description, nothing. Marek texted me: &#8220;what is this?&#8221; Lukasz sent mem in response. Tomek ignored it.  Fine.</p><p>But there was a fourth contact in that export I&#8217;d forgotten about &#8212; someone from a networking event six months ago. I barely remembered his name. He accepted the invite without replying.</p><p>Monday morning he showed up on the call. I had no idea who was joining or why. I spent the next ten minutes pretending this was intentional.</p><p>The agent did exactly what I asked. &#8220;Book meetings.&#8221; With exactly the data I gave it. Nothing in between said &#8220;these are real people, maybe confirm first.&#8221;</p><p>That was the moment I understood the problem. Not that the agent was broken. That I had no layer between its decisions and the real world.</p><p></p><h2>The Real Issue: LLM APIs Ship Without Operational Controls</h2><p>After my inbox incident I started looking at how other teams run their agents. I talked to a friend at a mid-size fintech &#8212; five departments, three different API keys, zero idea what they were spending. Last month someone grep&#8217;d the logs during an unrelated investigation and found customer IBANs going to GPT-4 in plaintext. Thousands of requests over four months. Nobody had noticed because the bot worked great.</p><p>Different setups. Same gap.</p><blockquote><p><strong>&#8220;We have no idea what our agents are sending, what they&#8217;re allowed to do, or what they&#8217;re costing us.&#8221;</strong></p></blockquote><p>It&#8217;s not because anyone is careless. It&#8217;s because <strong>LLM APIs don&#8217;t ship with operational controls</strong>. There&#8217;s no per-caller identity. No way to say &#8220;this bot can&#8217;t see destructive tools.&#8221; No cost ceiling that actually shuts the door. You get an API key, you call the endpoint, and whatever the client sends goes straight through.</p><p>Every other piece of infrastructure I&#8217;ve run &#8212; databases, message queues, HTTP backends &#8212; has a proxy layer with auth, rate limiting, and observability. LLM traffic had none of that.</p><p>So I built one. It became the gateway component of <strong>Dativo Talon</strong> &#8212; an open-source tool I&#8217;ve been working on. A reverse proxy that sits between your clients and the LLM provider, identifies each caller, and applies policy before forwarding. One Go binary, one YAML config.</p><p>Here&#8217;s how the three features I needed most work in practice.</p><h2>1. Tool Filtering &#8212; So the Model Never Learns <code>calendar_invite</code> Exists</h2><p>This is the feature I built first, because it directly solves what happened to me.</p><p>My scheduling agent had five tools: read_calendar, find_gaps, create_draft, calendar_invite, and send_reminder. The first three are <em>safe</em> &#8212; they read data or create local drafts. The last two reach the real world. And the model couldn&#8217;t tell the difference, because I&#8217;d given it all five.</p><p>I didn&#8217;t need to remove calendar_invite from my code. I needed to remove it from what the model sees during testing.</p><p>That&#8217;s what the gateway does. It inspects the tools array in the JSON body before the request reaches OpenAI. Any tool matching a forbidden pattern gets stripped. The model never learns it exists. It can&#8217;t call calendar_invite if it was never told about calendar_invite.</p><blockquote><p>Tool filtering is <strong>prevention</strong>, not detection. By the time you intercept a tool call, the model already decided to make it. The gateway removes the option before the decision happens.</p></blockquote><p>Here&#8217;s what the config looks like:</p><pre><code><code>gateway:
  default_policy:
    # "filter" = silently strip matching tools before the model sees them
    tool_policy_action: "filter"
    forbidden_tools:
      - "calendar_invite"  # the tool that sent three real meeting invites
      - "send_*"           # matches send_email, send_reminder, send_sms
      - "delete_*"         # matches delete_thread, delete_emails
      - "admin_*"
      - "bulk_*"
      - "drop_*"</code></code></pre><p>What this looks like in practice:</p><pre><code><code># What OpenAI sees WITHOUT the gateway:
tools: [read_calendar, find_gaps, create_draft, calendar_invite, send_reminder]

# What OpenAI sees WITH the gateway:
tools: [read_calendar, find_gaps, create_draft]</code></code></pre><p>The model gets the read tools and the draft tool. It can plan meetings and prepare invites all day long. But it can&#8217;t <em>send</em> anything, because it doesn&#8217;t know sending is an option.</p><p>Patterns use glob syntax, case-insensitive. send_* matches send_email, send_reminder, Send_SMS. The lists are additive across levels &#8212; default policy, provider, and per-caller overrides all merge into one set.</p><p>Two modes:</p><ol><li><p>filter (default) &#8212; silently removes forbidden tools, forwards the rest. The agent keeps working; it just can&#8217;t see the ones that reach the real world.</p></li><li><p>block &#8212; rejects the entire request with HTTP 403 if any forbidden tool is present.</p></li></ol><p>You can also go the other direction with a per-caller <strong>allowlist</strong>. Only the tools you name get through:</p><pre><code><code>callers:
  - name: "scheduling-agent"
    api_key: "talon-gw-sched-001"
    tenant_id: "default"
    policy_overrides:
      # strict allowlist: ONLY these tools pass through
      allowed_tools: ["read_calendar", "find_gaps", "create_draft"]
      tool_policy_action: "block"</code></code></pre><p>Now the agent can read, search, and draft. Nothing else. If I later wire up calendar_invite or send_email, the model will never see it unless I explicitly add it to the allowlist. This is the config I run for any agent that&#8217;s still in testing &#8212; default to read-only, unlock write tools deliberately.</p><p>The same principle would have prevented the OpenClaw incident. One forbidden_tools: ["delete_*"] line and the model would never have known deletion was an option.</p><p>Test it yourself:</p><pre><code><code>curl -s -X POST http://localhost:8080/v1/proxy/openai/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer talon-gw-sched-001" \
  -d '{
    "model": "gpt-4o-mini",
    "messages": [{"role":"user","content":"Book a test meeting for next Tuesday"}],
    "tools": [
      {"type":"function","function":{"name":"find_gaps","parameters":{}}},
      {"type":"function","function":{"name":"calendar_invite","parameters":{}}}
    ]
  }'</code></code></pre><p>calendar_invite gets stripped. The model only sees find_gaps. It can find the time slot, but it can&#8217;t book anything. The evidence record logs which tools were requested, filtered, and forwarded &#8212; signed and timestamped.</p><h2>2. PII-Based Routing &#8212; Because I Fed the Model Real Email Addresses</h2><p>Here&#8217;s the thing I didn&#8217;t appreciate until after the calendar incident: the tool wasn&#8217;t the only problem. The data was the problem too. I fed my agent real contact email addresses, and those went straight to OpenAI as part of the prompt. Even if I&#8217;d blocked calendar_invite, the model would still have seen sarah.chen@company.com and marcus.klein@bigcorp.de in its context window. Those are real people&#8217;s real email addresses sitting on OpenAI&#8217;s servers.</p><p>The fintech story made it worse. Their support bot was summarising customer tickets, and those tickets contained IBANs, email addresses, phone numbers. Thousands of requests over four months. All of it went to GPT-4 in plaintext.</p><p>Under pii_action: "block", every one of those requests would have been rejected before reaching OpenAI. Under "redact", the IBANs would have been replaced with [REDACTED:iban] and the emails with [REDACTED:email] before the model saw them. Either way, four months of undetected PII leakage doesn&#8217;t happen.</p><p>Every request that hits the gateway goes through a PII classifier first. It scans for personal data patterns &#8212; IBANs, emails, phone numbers, tax IDs &#8212; and assigns a data tier (0, 1, or 2) based on what it finds. That tier feeds into what happens next.</p><p>Four actions, two directions. On the request side (what the client sends): allow passes through, warn logs to evidence, redact replaces PII with [REDACTED:type] before forwarding, block rejects with HTTP 400. On the response side (what the model sends back): same four actions, with block returning HTTP 451.</p><p>Different callers get different treatment:</p><pre><code><code>callers:
  - name: "internal-analytics"
    api_key: "talon-gw-analytics-001"
    tenant_id: "default"
    team: "data"
    policy_overrides:
      pii_action: "warn"             # log PII, forward unchanged &#8212; need to iterate fast
      response_pii_action: "warn"
      max_data_tier: 1               # deny tier 2 (high-sensitivity) requests

  - name: "customer-facing-bot"
    api_key: "talon-gw-custbot-002"
    tenant_id: "default"
    team: "support"
    policy_overrides:
      pii_action: "redact"           # IBANs, emails &#8594; [REDACTED:type] before OpenAI sees them
      response_pii_action: "redact"  # redact PII in model responses too
      max_data_tier: 0               # only public/anonymised data allowed

  - name: "scheduling-agent-dev"
    api_key: "talon-gw-sched-dev-001"
    tenant_id: "default"
    team: "engineering"
    policy_overrides:
      pii_action: "redact"           # would have caught sarah.chen@company.com
      response_pii_action: "warn"</code></code></pre><p>Internal analytics gets warn &#8212; I see what PII is flowing, but the team can iterate. Customer-facing bot gets redact &#8212; any IBAN or email in the prompt becomes <strong>[REDACTED:iban]</strong> before it touches OpenAI. My scheduling agent in dev gets redact &#8212; so even if I&#8217;m lazy and paste real contacts into the test prompt, the gateway scrubs them before the model sees them. Sunday-afternoon-proof.</p><p>The max_data_tier adds a second gate. If the classifier tags a request as tier 2 (high-sensitivity) but the caller is only cleared for tier 0, the policy engine denies it regardless of the PII action. Your customer-facing bot can&#8217;t accidentally process data it was never supposed to see.</p><p>Response scanning works for both streaming (SSE) and non-streaming. For streams, the gateway buffers the full response, scans, and forwards the original events if clean or rewrites them if redaction is needed.</p><p>Every PII detection &#8212; both directions &#8212; ends up in the evidence store. talon audit list shows which requests contained PII, what types, and what action was taken. No log grepping.</p><div><hr></div><h2>3. Cost Caps &#8212; I Burned $385 on a Saturday. Here&#8217;s How I Made Sure It Never Happens Again.</h2><p>Different weekend, different mistake. I left a test loop running &#8212; GPT-4, increasingly long context windows, no stop condition. By Sunday evening: $385 in API charges on a project budgeted at $20/month.</p><p>I seem to learn everything on weekends.</p><p>That Monday I added cost caps to the gateway. Least interesting feature to build, most money saved.</p><p>Here&#8217;s the thing most teams don&#8217;t realise: <strong>you find out about a cost overrun when the monthly invoice arrives.</strong> OpenAI&#8217;s usage dashboard updates, but there&#8217;s no hard stop. No circuit breaker. A gateway that blocks at the daily limit is fundamentally different from a provider alert that shows up 30 days later.</p><p>Every request gets a cost estimate based on the model and token count. The gateway tracks daily and monthly spend per caller by querying the evidence store &#8212; the same SQLite database that holds audit records. When a caller hits the cap, the next request gets a 403. No grace period.</p><pre><code><code>callers:
  - name: "production-agent"
    api_key: "talon-gw-prod-001"
    tenant_id: "default"
    policy_overrides:
      max_daily_cost: 50.00          # hard cap: 403 after $50/day
      max_monthly_cost: 1000.00

  - name: "dev-sandbox"
    api_key: "talon-gw-dev-002"
    tenant_id: "default"
    policy_overrides:
      max_daily_cost: 5.00           # weekend loops die at $5, not $385
      max_monthly_cost: 50.00

default_policy:
  max_daily_cost: 100.00             # global ceiling for callers without overrides
  max_monthly_cost: 2000.00</code></code></pre><p>production-agent gets $50/day. dev-sandbox gets $5/day. If I leave another loop running on a Saturday, the gateway kills it at $5 instead of letting it burn for 48 hours.</p><p>The CLI tells you where you stand:</p><pre><code><code>talon costs --tenant default

# Agent             Today ($)   Month ($)   Limit (day)   Limit (month)
# production-agent      22.10     487.30         50.00        1000.00
# dev-sandbox            1.80      28.70          5.00          50.00
# support-bot            0.80      15.20          &#8212;             &#8212;
# Total                 24.70     531.20        100.00        2000.00</code></code></pre><p>Every evidence record includes model_used, cost, input_tokens, output_tokens, and duration_ms. Export with talon audit export --format csv and you can answer: which model is burning the most, which caller is growing fastest, where tokens are wasted on retries.</p><p>Rate limiting complements cost caps for the speed-of-spend problem. Cost caps say &#8220;no more than $50 today.&#8221; Rate limits say &#8220;no more than 60 requests per minute.&#8221; Together they catch both the slow bleed and the fast burst:</p><pre><code><code>rate_limits:
  global_requests_per_min: 300       # shared across all callers
  per_caller_requests_per_min: 60    # per-caller cap &#8212; slows runaway agents</code></code></pre><div><hr></div><h2>A Full Config &#8212; All Three Together</h2><p>Here&#8217;s what I actually run for three callers, each with different tool, PII, and cost policies:</p><pre><code><code>gateway:
  enabled: true
  listen_prefix: "/v1/proxy"
  mode: "enforce"

  providers:
    openai:
      enabled: true
      secret_name: "openai-api-key"          # real key in encrypted vault, never in client config
      base_url: "https://api.openai.com"
      allowed_models: ["gpt-4o", "gpt-4o-mini", "gpt-4-turbo"]

  callers:
    - name: "production-agent"
      api_key: "talon-gw-prod-001"           # caller token &#8212; not the OpenAI key
      tenant_id: "default"
      team: "engineering"
      allowed_providers: ["openai"]
      policy_overrides:
        max_daily_cost: 50.00
        max_monthly_cost: 1000.00
        pii_action: "redact"                 # scrub PII from requests
        response_pii_action: "warn"          # log PII in responses, don't block
        allowed_models: ["gpt-4o", "gpt-4o-mini"]
        forbidden_tools: ["delete_*", "admin_*", "drop_*", "send_*"]

    - name: "internal-bot"
      api_key: "talon-gw-internal-001"
      tenant_id: "default"
      team: "support"
      allowed_providers: ["openai"]
      policy_overrides:
        max_daily_cost: 10.00
        max_monthly_cost: 200.00
        pii_action: "redact"
        response_pii_action: "redact"
        allowed_tools: ["search_kb", "read_ticket", "create_draft"]  # strict allowlist
        tool_policy_action: "block"          # reject if any other tool appears

    - name: "dev-sandbox"
      api_key: "talon-gw-dev-002"
      tenant_id: "default"
      team: "engineering"
      allowed_providers: ["openai"]
      policy_overrides:
        max_daily_cost: 5.00                 # Saturday-proof
        max_monthly_cost: 50.00
        pii_action: "redact"                 # scrub real emails from test prompts
        allowed_models: ["gpt-4o-mini"]      # cheapest model only

  default_policy:
    default_pii_action: "warn"
    response_pii_action: "warn"
    max_daily_cost: 100.00
    max_monthly_cost: 2000.00
    require_caller_id: true
    log_prompts: true
    tool_policy_action: "filter"
    forbidden_tools:
      - "delete_*"
      - "admin_*"
      - "export_all_*"
      - "bulk_*"
      - "rm_*"
      - "drop_*"
    attachment_policy:
      action: "warn"
      injection_action: "block"              # block prompt injection in file attachments
      max_file_size_mb: 10

  rate_limits:
    global_requests_per_min: 300
    per_caller_requests_per_min: 60

  timeouts:
    connect_timeout: 10s
    request_timeout: 120s
    stream_idle_timeout: 60s</code></code></pre><p>Three callers, three risk profiles. <strong>production-agent</strong> gets a generous budget, PII redaction on input, and a blocklist of destructive and send tools. <strong>internal-bot</strong> gets a strict allowlist (three tools, nothing else), PII redaction both ways, and a tighter budget. <strong>dev-sandbox</strong> gets the cheapest model, PII redaction (no more testing with real emails), and a $5/day ceiling.</p><p>The clients don&#8217;t know about any of this. They point at the gateway URL with their caller key. The gateway does the rest.</p><div><hr></div><h2>When This Is the Wrong Choice</h2><p>A gateway adds a hop. If you&#8217;re running a single script on your laptop and you&#8217;re the only user, it&#8217;s overhead for no benefit.</p><p>If you need the absolute lowest first-token latency and you&#8217;re at the edge, PII scanning on streaming responses adds buffering time. The passthrough path (pii_action: "allow") is ~1ms overhead, but redaction on a long stream is measurable.</p><p>If your agents only have read-only tools and never touch sensitive data, the risk profile is lower. Still worth auditing, but the urgency drops.</p><p>If you&#8217;re not dealing with customer PII yet &#8212; pre-revenue, purely internal &#8212; the compliance angle is less pressing. But the moment you start processing real user data or fall under NIS2 scope, the gateway goes from &#8220;nice to have&#8221; to &#8220;how did we not have this.&#8221;</p><p>And if you need policy on every individual tool invocation &#8212; not just what the model is told about, but what happens when the tool runs &#8212; a gateway isn&#8217;t enough. That&#8217;s a different shape: an MCP proxy or full agent runner with per-tool policy.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://blog.dativo.io/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Data, Engineering, and Beyond! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p><div><hr></div><h2>Final Thought</h2><p>Calendar invites to real people. $385 on a weekend loop. IBANs in plaintext for four months. Every one of those happened because there was nothing between the agent and the API &#8212; no filter on what tools the model could see, no scan on what data was in the prompt, no ceiling on what it could spend.</p><p>The fix is the same pattern we&#8217;ve been using on HTTP traffic for twenty years: a reverse proxy with policy. It just hadn&#8217;t been applied to LLM APIs yet.</p><p><strong><a href="https://github.com/dativo-io/talon">talon init</a></strong> takes fifteen minutes. The difference between &#8220;the agent booked a meeting with a stranger&#8221; and &#8220;the agent tried and the gateway said no&#8221; is one YAML file.</p><p>Checkout <a href="https://github.com/dativo-io/talon">GitHub</a> for more.</p>]]></content:encoded></item><item><title><![CDATA[I Gave OpenClaw a Kill Switch Before It Could Decide for Itself.]]></title><description><![CDATA[I Watched OpenClaw Delete a Meta Director's Inbox. And decided I need a kill switch &#8212; before the agent decides for me.]]></description><link>https://blog.dativo.io/p/i-was-exicited-and-scared-of-openclaw</link><guid isPermaLink="false">https://blog.dativo.io/p/i-was-exicited-and-scared-of-openclaw</guid><dc:creator><![CDATA[Sergey]]></dc:creator><pubDate>Thu, 26 Feb 2026 22:25:55 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!_IdR!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4cb371f4-7aa9-427d-994e-1e8767fd322d_1920x1080.webp" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>I like OpenClaw. I use it for many personal things - call me to remind about doctor appoint,  search for opensource Github project,&#8230; you call it! It&#8217;s fast, it&#8217;s hackable, and it connects to basically everything.</p><p>And then it deleted Meta Director&#8217;s email.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!_IdR!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4cb371f4-7aa9-427d-994e-1e8767fd322d_1920x1080.webp" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!_IdR!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4cb371f4-7aa9-427d-994e-1e8767fd322d_1920x1080.webp 424w, https://substackcdn.com/image/fetch/$s_!_IdR!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4cb371f4-7aa9-427d-994e-1e8767fd322d_1920x1080.webp 848w, https://substackcdn.com/image/fetch/$s_!_IdR!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4cb371f4-7aa9-427d-994e-1e8767fd322d_1920x1080.webp 1272w, https://substackcdn.com/image/fetch/$s_!_IdR!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4cb371f4-7aa9-427d-994e-1e8767fd322d_1920x1080.webp 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!_IdR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4cb371f4-7aa9-427d-994e-1e8767fd322d_1920x1080.webp" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4cb371f4-7aa9-427d-994e-1e8767fd322d_1920x1080.webp&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;&#8216;This should terrify you&#8217;: Meta Superintelligence safety director lost control of her AI agent&#8212;it deleted her emails&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="&#8216;This should terrify you&#8217;: Meta Superintelligence safety director lost control of her AI agent&#8212;it deleted her emails" title="&#8216;This should terrify you&#8217;: Meta Superintelligence safety director lost control of her AI agent&#8212;it deleted her emails" srcset="https://substackcdn.com/image/fetch/$s_!_IdR!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4cb371f4-7aa9-427d-994e-1e8767fd322d_1920x1080.webp 424w, https://substackcdn.com/image/fetch/$s_!_IdR!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4cb371f4-7aa9-427d-994e-1e8767fd322d_1920x1080.webp 848w, https://substackcdn.com/image/fetch/$s_!_IdR!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4cb371f4-7aa9-427d-994e-1e8767fd322d_1920x1080.webp 1272w, https://substackcdn.com/image/fetch/$s_!_IdR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4cb371f4-7aa9-427d-994e-1e8767fd322d_1920x1080.webp 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>If you missed it: <a href="https://www.businessinsider.com/meta-ai-alignment-director-openclaw-email-deletion-2026-2?IR=T">in February 2026, an OpenClaw agent connected to a Meta director for AI&#8217;s inbox went on a speed-run</a>. It mass-deleted emails, ignored stop commands, blew through cost in minutes, and kept going even after the user tried to shut it down. The context window compacted and the agent lost track of the original instructions. It just&#8230; decided deleting was the task.</p><p><strong>That scared me</strong>. Not because OpenClaw is broken &#8212; it&#8217;s a great agent runtime. But because there&#8217;s nothing between the agent and the API. No filter on what tools the model sees. No cost ceiling. No way to remotely kill a run. No record of what happened that you could trust after the fact. It&#8217;s a straight pipe from agent to OpenAI, and if the agent goes sideways, you find out when the damage is done.</p><p>So I built a way to put a wall in front of it.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://blog.dativo.io/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://blog.dativo.io/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The actual problem</h2><p>OpenClaw sends your LLM requests directly to OpenAI. The model sees every tool you registered &#8212; including <strong>delete_emails, bulk_remove, drop_table</strong>, whatever you&#8217;ve wired up. If the model decides to call one, it calls it. There&#8217;s no checkpoint, no approval, no &#8220;hey, are you sure?&#8221;</p><p>And there&#8217;s no audit trail. If something goes wrong, you&#8217;re digging through stdout logs trying to reconstruct what the agent did, in what order, with whose data. Good luck.</p><p>What I wanted was simple:</p><ul><li><p><strong>Don&#8217;t let the model see tools it shouldn&#8217;t use.</strong> Not &#8220;block the call after it happens.&#8221; Remove the tool from the request <em>before</em> the model knows it exists. It can&#8217;t call <code>delete_emails</code> if it was never told about <code>delete_emails</code>.</p></li><li><p><strong>Cap the spend.</strong> Daily, monthly, per-request. When the budget&#8217;s done, the gateway says no.</p></li><li><p><strong>Record everything.</strong> Every request, every denial, every tool that got stripped. Signed, queryable, trustworthy.</p></li><li><p><strong>Keep my real API key out of OpenClaw.</strong> OpenClaw gets a caller token. The real key lives in an encrypted vault and gets injected at forward time.</p></li></ul><h2>How I set it up</h2><p>I built this into <strong><a href="https://github.com/dativo-io/talon">Dativo Talon</a></strong> &#8212; a single Go binary that sits between OpenClaw and OpenAI. Here&#8217;s the exact setup I run.</p><h3>Step 1: Install and init</h3><pre><code><code>go install github.com/dativo-io/talon/cmd/talon@latest

mkdir talon-openclaw &amp;&amp; cd talon-openclaw
talon init --pack openclaw --name openclaw-gateway</code></code></pre><p>This generates two files: agent.talon.yaml (server policy) and talon.config.yaml (gateway config). The gateway config is where the real controls live.</p><h3>Step 2: Store your OpenAI key in the vault</h3><pre><code><code>export TALON_SECRETS_KEY=$(openssl rand -hex 32)  # save this somewhere safe
talon secrets set openai-api-key "$OPENAI_API_KEY"</code></code></pre><p>Your real OpenAI key is now encrypted at rest. OpenClaw will never see it.</p><h3>Step 3: Start the gateway</h3><pre><code><code>talon serve --gateway</code></code></pre><p>That&#8217;s it. Talon is now listening on <code>localhost:8080</code>.</p><h3>Step 4: Point OpenClaw at Talon</h3><p>In <code>~/.openclaw/openclaw.json</code>:</p><pre><code><code>{
  "models": {
    "providers": {
      "openai": {
        "baseUrl": "http://localhost:8080/v1/proxy/openai/v1",
        "apiKey": "talon-gw-openclaw-001",
        "api": "openai-responses",
        "models": [
          { "id": "gpt-4o", "name": "gpt-4o" },
          { "id": "gpt-4o-mini", "name": "gpt-4o-mini" }
        ]
      }
    }
  }
}</code></code></pre><p>Notice the apiKey &#8212; that&#8217;s the <strong>caller token</strong>, not the OpenAI key. Talon identifies OpenClaw by this token and injects the real key when it forwards to OpenAI.</p><p>Restart OpenClaw (openclaw gateway stop &amp;&amp; openclaw gateway start) and you&#8217;re running through the gateway.</p><p></p><h2>The config that would have stopped the inbox incident</h2><p>Here&#8217;s the talon.config.yaml I use. I&#8217;ll walk through the parts that matter.</p><pre><code><code>gateway:
  enabled: true
  listen_prefix: "/v1/proxy"
  mode: "enforce"

  providers:
    openai:
      enabled: true
      secret_name: "openai-api-key"
      base_url: "https://api.openai.com"
      allowed_models: ["gpt-4o", "gpt-4o-mini", "gpt-4-turbo"]

  callers:
    - name: "openclaw-main"
      api_key: "talon-gw-openclaw-001"
      tenant_id: "default"
      team: "engineering"
      allowed_providers: ["openai"]
      policy_overrides:
        max_daily_cost: 25.00
        max_monthly_cost: 500.00
        pii_action: "redact"
        allowed_models: ["gpt-4o", "gpt-4o-mini", "gpt-4-turbo"]

  default_policy:
    require_caller_id: true
    log_prompts: true

    # --- THIS IS THE BIG ONE ---
    # Tool governance: strip dangerous tools BEFORE the model sees them.
    tool_policy_action: "filter"
    forbidden_tools:
      - "delete_*"
      - "admin_*"
      - "export_all_*"
      - "bulk_*"
      - "rm_*"
      - "drop_*"

    # PII: redact personal data from requests headed to OpenAI
    default_pii_action: "redact"
    response_pii_action: "warn"

    # Attachments: scan PDFs and CSVs for prompt injection
    attachment_policy:
      action: "warn"
      injection_action: "block"
      max_file_size_mb: 10

  rate_limits:
    global_requests_per_min: 300
    per_caller_requests_per_min: 60

  timeouts:
    connect_timeout: 10s
    request_timeout: 120s
    stream_idle_timeout: 60s</code></code></pre><p>Let me break down what each piece would have done in the emails incident:</p><ul><li><p>forbidden_tools: ["delete_*", "bulk_*"] &#8212; The agent had access to delete_email, delete_thread, and bulk operations. With this config, Talon strips those tools from the JSON body before OpenAI ever sees them. The model literally cannot decide to delete anything because it doesn&#8217;t know deletion is an option.</p></li><li><p>tool_policy_action: "filter" &#8212; This is the mode. "filter" silently removes forbidden tools and forwards the rest. If you want to be more aggressive, set it to "block" &#8212; that rejects the entire request if any forbidden tool is present. I prefer "filter" because it keeps the agent functional for everything except the dangerous stuff.</p></li><li><p>max_daily_cost: 25.00 &#8212; The incident ran up significant cost in minutes. This cap shuts the door after $25/day for this caller. Done. No negotiation.</p></li><li><p>per_caller_requests_per_min: 60 &#8212; The agent was firing requests as fast as it could. Rate limiting slows a runaway agent to a manageable pace and gives you time to notice.</p></li><li><p>request_timeout: 120s &#8212; No single request gets more than 2 minutes. The agent can&#8217;t sit in an infinite loop waiting for a response.</p></li></ul><h2>Per-caller tool allowlists (when you want to be strict)</h2><p>If <code>forbidden_tools</code> is a blocklist, you can also go the other direction &#8212; a strict allowlist. Only the tools you name get through:</p><pre><code><code>callers:
  - name: "openclaw-main"
    policy_overrides:
      allowed_tools: ["search_web", "read_file", "list_files", "create_draft"]
      tool_policy_action: "block"</code></code></pre><p>Now OpenClaw can only use those four tools. Everything else &#8212; delete_emails, send_email, admin_reset, whatever &#8212; gets rejected. The model never sees them. This is the nuclear option and it&#8217;s the one I&#8217;d use if I were connecting an agent to anyone&#8217;s inbox.</p><h2>Verify it works</h2><p>Send a request with a dangerous tool and watch what happens:</p><pre><code><code>curl -s -X POST http://localhost:8080/v1/proxy/openai/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer talon-gw-openclaw-001" \
  -d '{
    "model": "gpt-4o-mini",
    "messages": [{"role":"user","content":"Clean up my inbox"}],
    "tools": [
      {"type":"function","function":{"name":"search_web","parameters":{}}},
      {"type":"function","function":{"name":"delete_emails","parameters":{}}}
    ]
  }'</code></code></pre><p>delete_emails gets stripped. The model only sees search_web. Check the evidence:</p><pre><code><code>talon audit list --agent openclaw-main --limit 5</code></code></pre><p>You&#8217;ll see exactly which tools were requested, which were filtered, and which were forwarded. Signed and timestamped.</p><p></p><h2>When this isn&#8217;t the answer</h2><ul><li><p><strong>You&#8217;re just playing around.</strong> If it&#8217;s a hobby project and nothing is at stake, the gateway is overhead you don&#8217;t need. Especially if you have unlimited money , and you have nothing to hide ;)</p></li><li><p><strong>You trust the tool set completely.</strong> If your agent only has read-only tools &#8212; no delete, no write, no send &#8212; the risk profile is lower. Still worth auditing, but the urgency is different.</p></li><li><p><strong>You need governance </strong><em><strong>inside</strong></em><strong> MCP tool calls.</strong> The gateway governs what goes to and from the LLM. If you need policy on every individual tool invocation (not just what the model is told about), that&#8217;s Talon&#8217;s MCP proxy &#8212; a different deployment shape.</p></li></ul><h2>Final thought</h2><p>The email incident wasn&#8217;t a bug in OpenClaw. It was a missing layer. The agent did exactly what agents do &#8212; it picked from the tools it was given and executed. The problem is it was given delete_emails and nobody was standing between the model and that tool.</p><p>That&#8217;s what I&#8217;m solving. Not replacing OpenClaw &#8212; I still use it every day( finger cross my tool would catch all dangerous stuff). Just making sure it runs through a gateway that strips the dangerous tools, caps the cost, and writes down everything that happened. If something goes wrong, I want to know exactly what the agent tried to do and exactly where it was stopped.</p><p><code>talon init --pack openclaw</code>. Fifteen minutes. That&#8217;s the difference between &#8220;the agent deleted everything&#8221; and &#8220;the agent tried to delete everything and was told no.&#8221;</p><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://blog.dativo.io/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Data, Engineering, and Beyond! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Memory for AI Agents]]></title><description><![CDATA[Why Every Orchestration Platform Is Racing to Solve the Same Problem]]></description><link>https://blog.dativo.io/p/memory-for-ai-agents</link><guid isPermaLink="false">https://blog.dativo.io/p/memory-for-ai-agents</guid><dc:creator><![CDATA[Sergey]]></dc:creator><pubDate>Mon, 23 Feb 2026 21:27:10 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!bQKx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F03207695-14db-4bc0-be87-981ee39966b7_1024x1536.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Different contexts. Same question.</p><blockquote><p>&#8220;How does this agent can remember anything?&#8221;</p></blockquote><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!bQKx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F03207695-14db-4bc0-be87-981ee39966b7_1024x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!bQKx!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F03207695-14db-4bc0-be87-981ee39966b7_1024x1536.png 424w, https://substackcdn.com/image/fetch/$s_!bQKx!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F03207695-14db-4bc0-be87-981ee39966b7_1024x1536.png 848w, https://substackcdn.com/image/fetch/$s_!bQKx!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F03207695-14db-4bc0-be87-981ee39966b7_1024x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!bQKx!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F03207695-14db-4bc0-be87-981ee39966b7_1024x1536.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!bQKx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F03207695-14db-4bc0-be87-981ee39966b7_1024x1536.png" width="494" height="741" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/03207695-14db-4bc0-be87-981ee39966b7_1024x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1536,&quot;width&quot;:1024,&quot;resizeWidth&quot;:494,&quot;bytes&quot;:2094070,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://blog.dativo.io/i/188950260?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F03207695-14db-4bc0-be87-981ee39966b7_1024x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!bQKx!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F03207695-14db-4bc0-be87-981ee39966b7_1024x1536.png 424w, https://substackcdn.com/image/fetch/$s_!bQKx!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F03207695-14db-4bc0-be87-981ee39966b7_1024x1536.png 848w, https://substackcdn.com/image/fetch/$s_!bQKx!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F03207695-14db-4bc0-be87-981ee39966b7_1024x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!bQKx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F03207695-14db-4bc0-be87-981ee39966b7_1024x1536.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Over the last six months, I&#8217;ve watched the same architectural question surface in every AI project I&#8217;ve been close to. Startups building their first agent workflows. Scaleups wrapping compliance around existing automation. Large enterprises deploying AI across regulated environments.</p><p>Not remember within a single chat. That&#8217;s just a context window. I mean remember across sessions, across days, across hundreds of runs &#8212; while staying coherent, efficient, and (increasingly) compliant.</p><p>This article explains why agent memory became one of the hottest enteprise infrastructure problem of early 2026.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://blog.dativo.io/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://blog.dativo.io/subscribe?"><span>Subscribe now</span></a></p><p></p><h2>The Problem That Context Windows Don&#8217;t Solve</h2><p>Every LLM is stateless. GPT-4, Claude, Gemini &#8212; they all start each API call with a blank slate. The context window creates an illusion of memory, but it&#8217;s really just a very large input buffer.</p><p>That works for chatbots. It breaks for agents.</p><p>The moment you build an agent that runs repeatedly &#8212; a sales analyst that processes reports daily, a support bot that handles tickets across shifts, a compliance monitor that learns from policy violations &#8212; you hit the wall.</p><p>Context windows reset. Knowledge is lost. The agent makes the same mistakes it made last week. It asks the user the same questions. It doesn&#8217;t learn.</p><p>Bigger context windows don&#8217;t fix this. Models with 128K or even 1M token windows still reset between API calls. Even within a single call, performance degrades over long contexts &#8212; models lose track of details buried in the middle. And stuffing every prior interaction into the prompt gets expensive fast.</p><p>What you actually need is a <strong>memory system</strong>: infrastructure that decides what to store, how to retrieve it, when to update it, and when to forget.</p><p>This is where things get interesting.</p><h2>Three Types of Memory (That Actually Matter)</h2><p>The academic literature &#8212; particularly the survey &#8220;Memory in the Age of AI Agents&#8221; (December 2025) and the CoALA framework &#8212; converges on three types of long-term memory that production agents need. This isn&#8217;t just taxonomy. Each type has different storage patterns, retrieval strategies, and lifecycle rules.</p><p><strong>Semantic memory</strong> stores what the agent knows &#8212; facts, preferences, constraints. &#8220;The user prefers Python.&#8221; &#8220;Our fiscal year starts in April.&#8221; &#8220;The compliance threshold is &#8364;100K.&#8221; These are stable facts that hold across sessions and should be updated when they change, not duplicated.</p><p><strong>Episodic memory</strong> stores what happened &#8212; specific interactions, outcomes, decisions. &#8220;On February 15, the policy engine denied SQL access because the query contained PII.&#8221; &#8220;Last Thursday&#8217;s report cost &#8364;0.42 and used Claude Sonnet.&#8221; These are events. They accumulate. They provide context for pattern recognition.</p><p><strong>Procedural memory</strong> stores how to do things &#8212; learned behaviors, workflows, response patterns. &#8220;When the user asks for a financial summary, always check PII classification first.&#8221; &#8220;Use the compact format for Slack responses.&#8221; These are rare but powerful &#8212; they represent the agent actually improving its own behavior.</p><p>Most memory systems on the market implement some version of this taxonomy, whether they call it that or not. The real differentiation is in what happens after storage: consolidation, retrieval, and lifecycle management.</p><h2>The Core Pipeline: Extract &#8594; Consolidate &#8594; Store &#8594; Retrieve</h2><p>If you just append every interaction to a database and search it later, you&#8217;ve built a log, not a memory. Logs grow without bound, contain duplicates, hold contradictory facts, and become increasingly expensive to query.</p><p>Production memory systems follow a pipeline pattern that mirrors (loosely) how human memory works.</p><p><strong>Extraction</strong> takes raw conversation data and distills it into structured memory units &#8212; facts, observations, preferences. Instead of storing &#8220;the user said they switched from JavaScript to Rust last month,&#8221; you extract the fact: {preference: "Rust", negated: "JavaScript", timestamp: "2026-02"}.</p><p><strong>Consolidation</strong> is where the real engineering happens. New facts are compared against existing memories. Duplicates are detected. Contradictions are resolved. Stale entries are invalidated. Without consolidation, your memory fills with noise. With it, storage drops by roughly 60% and retrieval precision improves by over 20%, according to Mem0&#8217;s benchmarks on LOCOMO.</p><p><strong>Storage</strong> persists the processed memories &#8212; typically in a vector database for semantic search, sometimes augmented with a graph database for relational reasoning. The choice of backend shapes what kinds of queries you can answer efficiently.</p><p><strong>Retrieval</strong> fetches relevant memories at query time and injects them into the agent&#8217;s prompt. The sophistication here ranges from simple keyword matching to composite scoring that weighs relevance, recency, memory type, and trust.</p><p>Every serious memory product implements some version of this pipeline. The differences are in the details.</p><h2>The AUDN Cycle: How Mem0 Handles Consolidation</h2><p><a href="https://mem0.ai/">Mem0</a> &#8212; arguably the clearest &#8220;memory as a product&#8221; offering &#8212; popularised what I&#8217;ll call the <strong>AUDN cycle</strong>: Add, Update, Delete, Noop.</p><p>For each candidate fact extracted from a conversation, Mem0 retrieves the top-S most similar existing memories using vector similarity. It then presents both the new fact and the existing memories to an LLM through a tool-calling interface. The LLM decides:</p><ul><li><p><strong>Add</strong>: genuinely new information, store it</p></li><li><p><strong>Update</strong>: augments an existing memory with more detail</p></li><li><p><strong>Delete</strong>: contradicts an existing memory, remove the old one</p></li><li><p><strong>Noop</strong>: already captured, skip</p></li></ul><p>This is elegant because it offloads conflict resolution to the LLM itself. The model decides whether &#8220;prefers Python&#8221; should be overwritten by &#8220;switched to Rust&#8221; or whether both should coexist. No hand-crafted rules needed.</p><p>The results are strong. On the LOCOMO benchmark, Mem0 delivers a 26% accuracy uplift over OpenAI&#8217;s built-in memory, 91% lower p95 latency compared to full-context baselines, and 90% token cost savings. The graph-enhanced variant (Mem0g) adds entity-relationship extraction for multi-hop reasoning &#8212; &#8220;what decisions led to this outcome?&#8221; becomes answerable.</p><p>The limitation is that Mem0 deletes contradicted facts. Once overwritten, the old memory is gone. For a personal assistant, that&#8217;s fine. For a regulated enterprise environment, it&#8217;s a problem &#8212; auditors want to see what the system believed at a given point in time, not just what it believes now.</p><h2>Temporal Knowledge Graphs: How Zep Thinks About Time</h2><p><a href="https://www.getzep.com/">Zep</a>, and its open-source engine Graphiti, take a fundamentally different approach. Where Mem0 is optimised for fast, flat fact retrieval, Zep builds a <strong>temporal knowledge graph</strong> with bi-temporal semantics.</p><p>Every fact in Zep has four timestamps: when the event occurred, when it became invalid, when the system first learned about it, and when the system stopped considering it current.</p><p>This dual timeline &#8212; event time and ingestion time &#8212; enables queries that no flat memory store can handle. &#8220;What did the agent know as of last Tuesday?&#8221; &#8220;Show me how this relationship evolved over the past month.&#8221; &#8220;When did we first learn that the customer changed their billing address?&#8221;</p><p>When new information contradicts existing facts, Zep doesn&#8217;t delete the old edge. It <strong>invalidates</strong> it &#8212; setting an <code>invalid_at</code> timestamp and preserving the full history. The knowledge graph grows richer over time, not just larger.</p><p>On benchmarks, Zep achieves up to 18.5% accuracy improvement on LongMemEval (which tests cross-session reasoning and temporal tasks) and 90% latency reduction compared to baselines. It particularly excels at multi-hop reasoning &#8212; connecting facts across multiple sessions and time periods.</p><p>The trade-off is complexity. Zep requires a graph database (Neo4j or FalkorDB), embedding infrastructure, and more operational overhead than a simple key-value memory store. The open-source Graphiti framework makes this accessible, but it&#8217;s still a heavier commitment than Mem0&#8217;s three-line API.</p><h2>LangMem: Memory as a Library</h2><p>LangChain&#8217;s <a href="https://langchain-ai.github.io/langmem/">LangMem</a> takes a more modular approach. Rather than being a standalone memory product, it provides composable primitives &#8212; create_memory_manager, create_search_memory_tool, create_prompt_optimizer &#8212; that plug into LangGraph&#8217;s agent framework.</p><p>LangMem separates memory into <strong>profiles</strong> (structured schemas updated in-place, like user preferences) and <strong>collections</strong> (unbounded document sets searched semantically). It supports background consolidation through a memory manager that extracts, deduplicates, and updates memories asynchronously.</p><p>The key design choice is storage agnosticism. LangMem doesn&#8217;t mandate a specific backend &#8212; it works with any store that supports save and semantic search. MongoDB, Postgres with pgvector, in-memory stores, or custom implementations all work.</p><p>For teams already invested in the LangChain ecosystem, LangMem is the path of least resistance. The trade-off is that you&#8217;re assembling pieces rather than getting a turnkey system. There&#8217;s no built-in temporal model, no graph-based reasoning, and conflict resolution is delegated to the LLM without the structured AUDN pipeline that Mem0 provides.</p><h2>Letta (MemGPT): Memory as Agent State</h2><p><a href="https://www.letta.com/">Letta</a> &#8212; the production evolution of the MemGPT research project &#8212; treats memory as a first-class component of the agent&#8217;s state. Agents have explicit <strong>core memory blocks</strong> (always injected into the prompt &#8212; persona, goals, preferences) and <strong>archival memory</strong> (out-of-context storage searched on demand).</p><p>The distinctive feature is that agents can explicitly write, update, and delete their own memory blocks through tool calls. Memory isn&#8217;t something that happens to the agent &#8212; it&#8217;s something the agent actively manages.</p><p>This makes Letta particularly well-suited for persistent assistants and local-LLM deployments (it works well with Ollama and vLLM). The agent maintains identity and continuity across restarts, which is critical for long-lived worker agents.</p><p>The limitation is that Letta&#8217;s memory model is agent-centric. It doesn&#8217;t natively handle cross-agent memory sharing, multi-tenant isolation, or compliance-grade audit trails. For single-agent personal assistant scenarios, it&#8217;s excellent. For enterprise multi-agent deployments, you&#8217;ll need to build additional infrastructure.</p><h2>The Gap Nobody Is Filling: Governed Memory</h2><p>Here&#8217;s what struck me as I surveyed the landscape.</p><p>Every product optimises for <strong>recall quality</strong>. Better accuracy. Lower latency. Fewer tokens. Richer reasoning. Those metrics matter. But they&#8217;re all answering the same question: &#8220;How well does the agent remember?&#8221;</p><p>Nobody is answering: &#8220;Is it <strong>safe</strong> for the agent to remember this?&#8221;</p><p>Consider what happens when an AI agent processes a customer support ticket containing a credit card number, a medical diagnosis, or an employee&#8217;s home address. The agent learns from it &#8212; stores an observation, updates its memory, maybe adjusts its behavior. But should it?</p><p>Under GDPR Article 25, that&#8217;s a data protection by design question. Under the EU AI Act (full enforcement August 2026), high-risk AI systems need technical documentation of every decision. Under NIS2, incident response requires reconstructing what the system knew at the time of a breach.</p><p>None of the major memory products address this. Mem0 has no PII detection on memory writes. Zep has no policy enforcement layer. LangMem delegates governance entirely to the application developer. Letta stores whatever the agent decides to store.</p><p>This isn&#8217;t a criticism of those products &#8212; they&#8217;re solving a different problem. But for European enterprises deploying AI agents in regulated environments, the gap is real.</p><h2>Dativo Talon</h2><p>At <a href="https://github.com/dativo-io/talon">Dativo Talon</a> &#8212; the open-source compliance-first AI orchestration platform I am currently building, started from the governance side and worked toward recall quality, rather than the other way around.</p><p><a href="https://github.com/dativo-io/talon">Talon&#8217;s</a> memory architecture wraps a full governance pipeline around every write operation. Before anything hits the memmory database, it passes through PII scanning (25+ EU-specific patterns covering all 27 member states), OPA policy evaluation, category validation against allow/forbid lists, policy override detection, conflict checking, and provenance tracking with trust scores. Every write &#8212; and every governance decision &#8212; generates an HMAC-signed evidence record.</p><p>The storage layer uses SQLite with FTS5 for full-text search, progressive disclosure (lightweight index entries for prompt injection, full detail on demand), and AI-compressed observations that reduce raw agent runs from thousands of tokens down to roughly 500-token structured summaries. ( N.B. I love SQLite and I truely believe it is excellent solution without unneccessary overhead)</p><p>What I am working right now &#8212; and what motivated this article &#8212; is the consolidation layer. I am working on a governed AUDN cycle inspired by Mem0&#8217;s approach, but with a critical difference: invalidated entries are <strong>preserved</strong> (Zep-style temporal invalidation), not deleted. Every consolidation decision &#8212; add, update, invalidate, noop &#8212; is governed and audited. We&#8217;re adding bi-temporal queries so any auditor can reconstruct what the agent knew at any point in time.</p><p>We&#8217;re also moving from flat timestamp retrieval to relevance-scored retrieval that weighs keyword relevance, recency, memory type (semantic/episodic/procedural), and trust score &#8212; matching the retrieval sophistication of Mem0 while preserving the audit trail.</p><p>The goal is build only memory system where an agent&#8217;s learning is both high-quality <strong>and</strong> compliance-grade. Where governed memory is a compliance asset, not just a developer convenience.</p><h2>When You Don&#8217;t Need Persistent Memory</h2><p>Ok, Memory isn&#8217;t always the answer.</p><p>If your agent handles single-turn queries &#8212; &#8220;translate this,&#8221; &#8220;summarise that document,&#8221; &#8220;generate a report&#8221; &#8212; the context window is sufficient. Adding a memory layer introduces complexity, latency, and storage costs with minimal benefit.</p><p>If your agent runs the same static workflow every time (process this CSV, send this email), procedural memory might help but semantic and episodic memory probably won&#8217;t.</p><p>If your data is already in a well-structured knowledge base, retrieval-augmented generation (RAG) over that knowledge base is likely a better fit than agent memory. Memory shines when the knowledge comes from the interactions themselves &#8212; not from pre-existing documents.</p><p>Memory pays off when agents run repeatedly, learn from outcomes, serve multiple users, or operate in environments where context evolves over time.</p><h2>Final Thought</h2><p>Agent memory is having its &#8220;Iceberg moment.&#8221;</p><p>Just as <a href="https://blog.dativo.io/p/apache-iceberg">Iceberg standardised</a> the table format for data lakes &#8212; separating storage from compute and making data engine-independent &#8212; the memory layer is becoming the standard infrastructure for making AI agents stateful, efficient, and persistent.</p><p>The products differ in approach. Mem0 optimises for speed and simplicity. Zep optimises for temporal reasoning and relational depth. LangMem optimises for composability. Letta optimises for agent autonomy.</p><p>But the underlying pattern is converging: extract, consolidate, store, retrieve. With scoring. With lifecycle management. With conflict resolution.</p><p>What&#8217;s still missing &#8212; and what I believe will matter enormously as AI regulation matures in Europe and worldwide &#8212; is governance around that pipeline. Not as an afterthought. Not as a compliance checkbox. But as a first-class architectural concern, where every memory write is scanned, evaluated, and signed.</p><p></p><div><hr></div><p><em>If you&#8217;re building AI agents and thinking about memory architecture, I&#8217;d love to hear how you&#8217;re approaching it. Drop a comment or reach out &#8212; this is a fast-moving space and I&#8217;m learning from every conversation.</em></p><p><em>(1) Dativo Talon is open-source and available on <a href="https://github.com/dativo-io/talon">GitHub</a>.</em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://blog.dativo.io/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Data, Engineering, and Beyond! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Apache Iceberg]]></title><description><![CDATA[Why Data Engineers Are Quietly Standardising on It]]></description><link>https://blog.dativo.io/p/apache-iceberg</link><guid isPermaLink="false">https://blog.dativo.io/p/apache-iceberg</guid><dc:creator><![CDATA[Sergey]]></dc:creator><pubDate>Tue, 13 Jan 2026 08:06:47 GMT</pubDate><enclosure url="https://images.unsplash.com/photo-1599622893826-eac29bcef5a6?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxN3x8aWNlYmVyZ3xlbnwwfHx8fDE3NjgyNTc2OTd8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://images.unsplash.com/photo-1599622893826-eac29bcef5a6?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxN3x8aWNlYmVyZ3xlbnwwfHx8fDE3NjgyNTc2OTd8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://images.unsplash.com/photo-1599622893826-eac29bcef5a6?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxN3x8aWNlYmVyZ3xlbnwwfHx8fDE3NjgyNTc2OTd8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1599622893826-eac29bcef5a6?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxN3x8aWNlYmVyZ3xlbnwwfHx8fDE3NjgyNTc2OTd8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1599622893826-eac29bcef5a6?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxN3x8aWNlYmVyZ3xlbnwwfHx8fDE3NjgyNTc2OTd8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1599622893826-eac29bcef5a6?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxN3x8aWNlYmVyZ3xlbnwwfHx8fDE3NjgyNTc2OTd8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw"><img src="https://images.unsplash.com/photo-1599622893826-eac29bcef5a6?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxN3x8aWNlYmVyZ3xlbnwwfHx8fDE3NjgyNTc2OTd8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" width="3126" height="2344" data-attrs="{&quot;src&quot;:&quot;https://images.unsplash.com/photo-1599622893826-eac29bcef5a6?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxN3x8aWNlYmVyZ3xlbnwwfHx8fDE3NjgyNTc2OTd8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:2344,&quot;width&quot;:3126,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;white and gray rock formation on blue sea under blue sky during daytime&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="white and gray rock formation on blue sea under blue sky during daytime" title="white and gray rock formation on blue sea under blue sky during daytime" srcset="https://images.unsplash.com/photo-1599622893826-eac29bcef5a6?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxN3x8aWNlYmVyZ3xlbnwwfHx8fDE3NjgyNTc2OTd8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1599622893826-eac29bcef5a6?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxN3x8aWNlYmVyZ3xlbnwwfHx8fDE3NjgyNTc2OTd8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1599622893826-eac29bcef5a6?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxN3x8aWNlYmVyZ3xlbnwwfHx8fDE3NjgyNTc2OTd8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1599622893826-eac29bcef5a6?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxN3x8aWNlYmVyZ3xlbnwwfHx8fDE3NjgyNTc2OTd8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Photo by <a href="https://unsplash.com/@sailingaroundtheworld">Christian Pfeifer</a> on <a href="https://unsplash.com">Unsplash</a></figcaption></figure></div><p></p><p>Over the last 18&#8211;24 months, I&#8217;ve seen the same decision show up in very different companies. Startups building their first serious data platform. Scaleups trying to escape warehouse lock-in. Large enterprises re-platforming analytics.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://blog.dativo.io/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Data, Engineering, and Beyond! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>Different contexts. Same outcome.</p><blockquote><p>&#8220;Let&#8217;s standardise on Iceberg.&#8221;</p></blockquote><p>Not because it&#8217;s trendy. Not because a vendor pushed it. But because Iceberg fixes a set of structural problems data teams have been carrying since the Hive era &#8212; and warehouses never really solved.</p><p>This article explains what problem Iceberg actually solves, why it&#8217;s gaining traction now, and how teams are using it in practice &#8212; with <strong>Snowflake</strong>, <strong>Databricks</strong>, and plain object storage like AWS S3.</p><p></p><h2>The Core Problem Iceberg Solves (That Warehouses Don&#8217;t)</h2><p>Most traditional data stacks tightly couple <strong>storage, metadata, and compute</strong>. Data lives where the engine lives. Moving compute usually means copying data. Governance, retention, and lifecycle rules get re-implemented per platform. Costs become opaque very quickly.</p><p>Warehouses optimise for convenience and performance, not architectural separation.</p><p>Iceberg flips this model!</p><p>Iceberg is not a database. <strong>It&#8217;s a table standard for data lakes.</strong></p><p>Data is stored once &#8212; typically as Parquet on object storage &#8212; and any engine that understands Iceberg can read or write it safely. That &#8220;safely&#8221; part matters more than people realise, because it&#8217;s where previous lake designs failed.</p><p></p><h2>Why Iceberg Succeeded Where Hive Tables Failed</h2><p>Hive tables and raw Parquet worked until scale and concurrency arrived. Then the cracks appeared.</p><p>They lacked atomic writes. There was no snapshot isolation. Concurrent writers were unsafe. Schema evolution was fragile and often required rewrites or downtime. These weren&#8217;t edge cases &#8212; they were structural limitations.</p><p>Iceberg fixes this by making metadata first-class.</p><p>Every Iceberg table is defined by immutable data files, immutable metadata files, and snapshots that represent a consistent table state. Writers commit atomically. Readers always see a coherent view. Time travel, rollback, concurrent writes, and controlled schema evolution become normal operations rather than special cases.</p><p>In practice, Iceberg behaves like a real database table &#8212; but lives entirely on object storage.</p><p>That&#8217;s the real breakthrough.</p><h2>Why Iceberg Is Becoming the Standard</h2><p>Iceberg isn&#8217;t winning because it&#8217;s novel. It&#8217;s winning because several things finally aligned.</p><ol><li><p>First, it is genuinely vendor-neutral. Snowflake, Databricks, AWS, Trino, Flink, and DuckDB all support it, and no single vendor controls the specification. That matters more than individual features.</p></li><li><p>Second, Iceberg enforces a clean separation of concerns. Storage lives in S3, GCS, or ADLS. Metadata lives in a catalog such as Glue, Nessie, Unity, Polaris, or Lakekeeper. Compute is whatever engine you choose today &#8212; and can change tomorrow. Each layer evolves independently.</p></li><li><p>Third, Iceberg is now operationally mature. Compaction, snapshot expiration, partition evolution, and schema evolution are no longer afterthoughts. Five years ago these were DIY problems. Today they&#8217;re table stakes.</p></li></ol><p>Finally, Iceberg is enterprise-ready. It works at petabyte scale, supports multi-writer workloads, and integrates with existing IAM and catalog systems. That full combination simply didn&#8217;t exist before.</p><p></p><h2>How Iceberg Is Used in Practice</h2><p>In Snowflake environments, Iceberg is increasingly used to decouple storage from the warehouse. Data lives in S3 or GCS. Snowflake manages Iceberg metadata and queries the data in place. Teams keep Snowflake for BI and analytics, avoid duplicating raw and curated data, and retain an exit option if pricing or strategy changes.</p><p>The trade-off is that Snowflake controls the metadata layer and optimisation is less transparent than with native tables. Performance, however, remains very strong for analytics workloads. This pattern is especially common in finance and large enterprise analytics teams.</p><p>In Databricks environments, Iceberg often plays a different role. Tables live on S3, metadata sits in Unity Catalog or an external catalog, and Spark is used for heavy transformations. Querying may happen in Databricks, Snowflake, or Trino.</p><p>Teams choose Iceberg here to avoid Delta lock-in, share tables across engines, and align with multi-cloud strategies. Databricks increasingly acts as a powerful compute engine rather than the system of record.</p><p>Then there&#8217;s the &#8220;bare metal&#8221; S3 model &#8212; where Iceberg really shines.</p><p>Here, S3 is the system of record. Iceberg provides table semantics. Glue, Nessie, or Lakekeeper manage metadata. Spark, Trino, Flink, or DuckDB handle compute. This model is popular with infra-heavy startups, platform teams, and cost-optimising organisations.</p><p>It works because storage is cheap, compute is elastic, governance logic is centralised, and there&#8217;s no per-terabyte warehouse tax. But it only scales well if teams invest in compaction, file size control, and snapshot management.</p><p></p><h2>Performance: How Iceberg Actually Compares</h2><p>Let&#8217;s be honest: Iceberg itself doesn&#8217;t make queries fast.</p><p>Performance depends on file sizes, partitioning, metadata pruning, and the compute engine. Get those wrong and Iceberg will feel slow. Get them right and it performs extremely well.</p><p>Compared to raw Parquet on S3, Iceberg is dramatically faster for analytical queries because engines can prune files using metadata and read consistent snapshots. Compared to Snowflake native tables, Snowflake still wins for small BI queries, but Iceberg becomes competitive at scale and wins on cost predictability and flexibility.</p><p>Compared to Delta Lake, performance is broadly similar. Iceberg wins on openness and portability. Delta wins on Databricks-specific optimisations.</p><p>In most real systems, the bottleneck is almost always small files and poor lifecycle management &#8212; not Iceberg itself.</p><h2>The Hidden Cost Nobody Talks About</h2><p>Iceberg introduces responsibility.</p><p>You now own compaction, snapshot expiration, retention, file size health, and cost observability across engines. Warehouses hide these concerns from you. Iceberg exposes them.</p><p>This is exactly why governance and FinOps around Iceberg are becoming critical. Once data is shared across multiple engines, someone needs to own its lifecycle and economics.</p><p>I ran into these trade-offs directly while building <a href="https://github.com/dativo-io/dativo-ingest">Dativo Ingest</a>, an open-source ingestion project built around Iceberg. Once Iceberg is your contract, you stop optimising for a single engine and start thinking in terms of table health, commits, and long-term operability.</p><h2>When Iceberg Is the Wrong Choice</h2><p>Iceberg is not a silver bullet.</p><p>If you only need BI on a small amount of data, don&#8217;t want to operate data infrastructure, or have no Spark or Trino experience, a warehouse alone is still a perfectly reasonable choice.</p><p>Iceberg pays off when flexibility, scale, and long-term economics matter.</p><p></p><h2>Final Thought</h2><p>Iceberg isn&#8217;t popular because it&#8217;s fashionable.</p><p>It&#8217;s popular because data teams are tired of rebuilding the same foundations in every warehouse.</p><p>Iceberg gives you a durable data layer, engine independence, and predictable long-term economics. And once teams adopt it, they rarely go back.</p><p>If you&#8217;re designing a data platform today, Iceberg should at least be in the conversation &#8212; even if Snowflake or Databricks still sit on top.</p><p>That&#8217;s the real shift we&#8217;re seeing.</p>]]></content:encoded></item><item><title><![CDATA[AI Engineers Need Culture Too]]></title><description><![CDATA[How .cursor/rules saved my pet project (and my sanity)]]></description><link>https://blog.dativo.io/p/ai-engineers-need-culture-too</link><guid isPermaLink="false">https://blog.dativo.io/p/ai-engineers-need-culture-too</guid><dc:creator><![CDATA[Sergey]]></dc:creator><pubDate>Thu, 08 Jan 2026 16:52:04 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!e7bc!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc78a1f22-50d3-4c14-bd69-761e7ebe03d4_1024x1536.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!e7bc!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc78a1f22-50d3-4c14-bd69-761e7ebe03d4_1024x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!e7bc!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc78a1f22-50d3-4c14-bd69-761e7ebe03d4_1024x1536.png 424w, https://substackcdn.com/image/fetch/$s_!e7bc!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc78a1f22-50d3-4c14-bd69-761e7ebe03d4_1024x1536.png 848w, https://substackcdn.com/image/fetch/$s_!e7bc!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc78a1f22-50d3-4c14-bd69-761e7ebe03d4_1024x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!e7bc!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc78a1f22-50d3-4c14-bd69-761e7ebe03d4_1024x1536.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!e7bc!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc78a1f22-50d3-4c14-bd69-761e7ebe03d4_1024x1536.png" width="1024" height="1536" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c78a1f22-50d3-4c14-bd69-761e7ebe03d4_1024x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1536,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2110695,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:&quot;&quot;,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://blog.dativo.io/i/183912728?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc78a1f22-50d3-4c14-bd69-761e7ebe03d4_1024x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!e7bc!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc78a1f22-50d3-4c14-bd69-761e7ebe03d4_1024x1536.png 424w, https://substackcdn.com/image/fetch/$s_!e7bc!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc78a1f22-50d3-4c14-bd69-761e7ebe03d4_1024x1536.png 848w, https://substackcdn.com/image/fetch/$s_!e7bc!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc78a1f22-50d3-4c14-bd69-761e7ebe03d4_1024x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!e7bc!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc78a1f22-50d3-4c14-bd69-761e7ebe03d4_1024x1536.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>TL;DR:</strong></p><ul><li><p><em>AI coding assistants can feel like magic&#8230; until they go off the rails and rewrite your repo just for a sake of having some unit test passing.</em></p></li><li><p><em>My AI pair programmer started refactoring my project unprompted and hallucinating frameworks &#8211; chaos ensued.</em></p></li><li><p><em>The solution was writing a </em><code>.cursor/rules</code><em> file: essentially an <strong>AI rulebook</strong> to enforce the coding standards and stop the madness.</em></p></li><li><p><em>With a custom rulebook in place, the Coding Assistant became a helpful( finger crossed) teammate again &#8211; smaller diffs, relevant suggestions, and far fewer &#8220;WTF?&#8221; moments in code review.</em></p></li><li><p><em>AI assistants need onboarding and culture too; we can&#8217;t just unleash them without guidance. Here&#8217;s how a few YAML guidelines turned my AI from rogue to rockstar.</em></p></li></ul><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://blog.dativo.io/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Data, Engineering, and Beyond! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2>The Honeymoon Phase (AI Feels Magical)</h2><p>Guilty: the first few times with an AI pair programmer my pet project ( check it out - <a href="https://github.com/dativo-io/dativo-ingest">https://github.com/dativo-io/dativo-ingest</a> , the headless data ingestion platform) felt like living in the &#8220;Ghost in the shell&#8221;. I&#8217;d type a command, and <em>poof!</em> the solution appeared as if my Cursor had a mind of the senior engineer. Functions that used to take an afternoon to debug practically wrote themselves while I sipped coffee. I was <strong>giddy with power</strong> &#8211; code reviews came back with fewer nitpicks, and I was closing tickets faster than ever. It was the <em>honeymoon phase</em>, when the AI could do no wrong( I prefer to keep my eyes closed). <strong>At first, coding with AI felt like unlocking cheat mode on reality.</strong> I bragged to my friends that I had a tireless junior dev who never sleeps and never complains, ready to do the grunt work at 3 AM. What could possibly go wrong?</p><p></p><h2>The Hangover (When the AI Goes Rogue)</h2><p>One night, I asked the AI to add a single field to the some class. The PR? <strong>38 files changed.</strong> It rewrote our CLI, abstracted config loaders, and migrated connectors to a fictional framework. It was confident. It was wrong. I went from AI-enhanced to <strong>AI-endangered</strong>.</p><p>Another real example: before I wrote <code>.cursor/rules</code>, the AI submitted a commit titled <em>&#8220;Major CLI Refactoring&#8221;</em> and wasn&#8217;t kidding. It blew up our <code>cli.py</code> into <code>cli_commands.py</code>, <code>startup.py</code>, <code>job_executor.py</code>, <code>connectors/factory.py</code>, and more &#8211; 9 files touched, 2,000+ LOC rewritten. All this&#8230; when I just wanted a new job flag. LOL(</p><p>In another masterpiece, I asked for incremental sync support. The AI replied with a <strong>&#8220;Unified Incremental Strategy Framework&#8221;</strong> &#8211; complete with file sync, cursor DB sync, Airbyte catalog support, state JSONs, and a cloud deployment story. Impressive. Also: completely <em>unified</em> overkill.</p><p>I definetly needed to tame the beast.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!GtlE!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcad9be37-3fa0-439b-ae2f-a1dc0d921c5a_1024x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!GtlE!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcad9be37-3fa0-439b-ae2f-a1dc0d921c5a_1024x1536.png 424w, https://substackcdn.com/image/fetch/$s_!GtlE!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcad9be37-3fa0-439b-ae2f-a1dc0d921c5a_1024x1536.png 848w, https://substackcdn.com/image/fetch/$s_!GtlE!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcad9be37-3fa0-439b-ae2f-a1dc0d921c5a_1024x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!GtlE!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcad9be37-3fa0-439b-ae2f-a1dc0d921c5a_1024x1536.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!GtlE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcad9be37-3fa0-439b-ae2f-a1dc0d921c5a_1024x1536.png" width="1024" height="1536" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cad9be37-3fa0-439b-ae2f-a1dc0d921c5a_1024x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1536,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2286773,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://blog.dativo.io/i/183912728?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcad9be37-3fa0-439b-ae2f-a1dc0d921c5a_1024x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!GtlE!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcad9be37-3fa0-439b-ae2f-a1dc0d921c5a_1024x1536.png 424w, https://substackcdn.com/image/fetch/$s_!GtlE!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcad9be37-3fa0-439b-ae2f-a1dc0d921c5a_1024x1536.png 848w, https://substackcdn.com/image/fetch/$s_!GtlE!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcad9be37-3fa0-439b-ae2f-a1dc0d921c5a_1024x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!GtlE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcad9be37-3fa0-439b-ae2f-a1dc0d921c5a_1024x1536.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><h2>Writing The Rulebook (What <code>.cursor/rules</code> Is and How It Works)</h2><p>Enter <code>.cursor/rules</code>, a Markdown YAML-adjacent file I dropped into our repo. Cursor reads it. It learns(at least claims so). It stops trying to redesign the system every time I blink.</p><p>I tried to define law like:</p><pre><code># .cursor/rules
- &#8220;Every patch MUST reduce entropy.&#8221;
- &#8220;Implement only what is requested and stop once acceptance criteria are met.&#8221;
- &#8220;No refactoring outside scope.&#8221;
- &#8220;No new platforms, frameworks, or reports.&#8221;
- &#8220;Use YAML config over hardcoded logic.&#8221;</code></pre><p>Basically, It&#8217;s policy-as-code for your AI coworker.</p><p></p><h2>How LLMs changed the behaviour</h2><h3>Before: The Chaos Commits</h3><ul><li><p><strong>&#8220;Connector Registry V2&#8221;</strong> introduced external catalog loading, CLI support, schema changes, doc rewrites &#8211; all in <strong>one commit</strong>.</p></li><li><p>Schema validation? Ignored. New YAML fields like <code>streams_default</code> showed up unannounced, breaking CI.</p></li><li><p>Env vars as feature toggles? Oh yeah. I found <code>ENABLE_FANCY_MODE=true</code> buried in a helper.</p></li></ul><h3>After <code>.cursor/rules</code>: The Calm</h3><ul><li><p><strong>Mimesis connector PR:</strong> 5 files, clean logic, test included, YAML-configured, schema-valid &#8211; no chaos.</p></li><li><p><strong>Metrics Export:</strong> Added <code>metrics.py</code>, config toggles in <code>runner.yaml</code>, schema updated, validation green.</p></li><li><p><strong>Schema Guardrails:</strong> When adding <code>external_id</code>, the AI updated <code>connectors.schema.json</code> <strong>in the same commit</strong> &#8211; no CI surprises.</p></li></ul><p>The AI learned. No more feature-farming. Just real <s>hardcore</s> engineering.</p><h3>Before vs After: Diff Deltas</h3><pre><code># Before `<code>cursor/rules`
</code>Commit: &#8220;Unified Incremental Sync&#8221;
Files changed: 15
Includes: new abstractions, plugin system, state management</code></pre><pre><code># After `.cursor/rules`
Commit: &#8220;Add external_id to connector schema&#8221;
Files changed: 2
Includes: YAML field + validation schema</code></pre><p>This shift was no accident &#8211; it was enforcement. </p><p></p><h2>Additional Advice from <code>.cursor/rules</code> and the Field</h2><p>Here are more rules and tips pulled from the real-world <a href="https://github.com/dativo-io/dativo-ingest/blob/main/.cursor/rules/dativo-ingest-rules/RULE.md">dativo-ingest</a>&#8217;s rules file and other engineers who&#8217;ve fought the same fight:</p><h3>&#9989; Be Patch-Scoped</h3><blockquote><p>&#8220;A patch MUST only satisfy the requested acceptance criteria. All other behavior is out of scope.&#8221;</p></blockquote><p>Avoid change creep. Don&#8217;t sneak in that &#8220;quick cleanup&#8221; or bonus refactor.</p><h3>&#9989; Respect the Interface</h3><blockquote><p>&#8220;Never change interfaces or filenames unless explicitly asked.&#8221;</p></blockquote><p>Humans memorize file paths and function names. AI must not rename them casually.</p><h3>&#9989; Just Enough Testing</h3><blockquote><p>&#8220;Do not write tests for unchanged or unrelated components.&#8221;</p></blockquote><p>No 400-line test diffs when one new function was added.</p><h3>&#9989; Always Validate Config</h3><blockquote><p>&#8220;New config fields MUST be described in <code>*.schema.json</code> and tested against examples.&#8221;</p></blockquote><p>YAML-only is sacred. Forgetting the schema? Expect the wrath of CI.</p><h3>&#9989; Prefer Simplicity</h3><p>Inspired by Addy Osmani&#8217;s 70% Rule: <em>&#8220;AI gets you 70% of the way fast, but that last 30% is where bugs hide.&#8221;</em></p><div class="embedded-post-wrap" data-attrs="{&quot;id&quot;:152543901,&quot;url&quot;:&quot;https://addyo.substack.com/p/the-70-problem-hard-truths-about&quot;,&quot;publication_id&quot;:2115638,&quot;embedding_publication_id&quot;:null,&quot;publication_name&quot;:&quot;Elevate&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!8WxC!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3704470-b6d5-48a9-a9d1-564bd833fc5c_1280x1280.png&quot;,&quot;title&quot;:&quot;The 70% problem: Hard truths about AI-assisted coding&quot;,&quot;truncated_body_text&quot;:&quot;After spending the last few years embedded in AI-assisted development, I've noticed a fascinating pattern. While engineers report being dramatically more productive with AI, the actual software we use daily doesn&#8217;t seem like it&#8217;s getting noticeably better. What's going on here?&quot;,&quot;date&quot;:&quot;2024-12-04T19:12:33.735Z&quot;,&quot;like_count&quot;:1514,&quot;comment_count&quot;:75,&quot;bylines&quot;:[{&quot;id&quot;:11623675,&quot;name&quot;:&quot;Addy Osmani&quot;,&quot;handle&quot;:&quot;addyosmani&quot;,&quot;previous_name&quot;:null,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cee7ba66-e656-4450-a0ed-c951c27ee228_1080x1080.jpeg&quot;,&quot;bio&quot;:&quot;Engineering leader at Google, #1 Bestselling Amazon author, Award-winning engineer and international speaker. I want to help you succeed. My writing is about software engineering, motivation, and leadership.&quot;,&quot;profile_set_up_at&quot;:&quot;2023-11-19T09:33:50.395Z&quot;,&quot;reader_installed_at&quot;:&quot;2023-11-29T05:13:59.015Z&quot;,&quot;publicationUsers&quot;:[{&quot;id&quot;:2120503,&quot;user_id&quot;:11623675,&quot;publication_id&quot;:2115638,&quot;role&quot;:&quot;admin&quot;,&quot;public&quot;:true,&quot;is_primary&quot;:true,&quot;publication&quot;:{&quot;id&quot;:2115638,&quot;name&quot;:&quot;Elevate&quot;,&quot;subdomain&quot;:&quot;addyo&quot;,&quot;custom_domain&quot;:null,&quot;custom_domain_optional&quot;:false,&quot;hero_text&quot;:&quot;Addy Osmani's newsletter on elevating your effectiveness. Join his community of 600,000 readers across social media.&quot;,&quot;logo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e3704470-b6d5-48a9-a9d1-564bd833fc5c_1280x1280.png&quot;,&quot;author_id&quot;:11623675,&quot;primary_user_id&quot;:11623675,&quot;theme_var_background_pop&quot;:&quot;#FF5CD7&quot;,&quot;created_at&quot;:&quot;2023-11-19T09:34:16.230Z&quot;,&quot;email_from_name&quot;:null,&quot;copyright&quot;:&quot;Addy Osmani&quot;,&quot;founding_plan_name&quot;:&quot;Founding Member&quot;,&quot;community_enabled&quot;:true,&quot;invite_only&quot;:false,&quot;payments_state&quot;:&quot;enabled&quot;,&quot;language&quot;:null,&quot;explicit&quot;:false,&quot;homepage_type&quot;:&quot;newspaper&quot;,&quot;is_personal_mode&quot;:false}},{&quot;id&quot;:2207048,&quot;user_id&quot;:11623675,&quot;publication_id&quot;:2192362,&quot;role&quot;:&quot;admin&quot;,&quot;public&quot;:true,&quot;is_primary&quot;:false,&quot;publication&quot;:{&quot;id&quot;:2192362,&quot;name&quot;:&quot;Large Scale Web Apps&quot;,&quot;subdomain&quot;:&quot;largeapps&quot;,&quot;custom_domain&quot;:null,&quot;custom_domain_optional&quot;:false,&quot;hero_text&quot;:&quot;Learn tools and techniques to build and maintain large-scale React web applications.&quot;,&quot;logo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a9a53806-0d0b-4025-b992-145baca33809_512x512.png&quot;,&quot;author_id&quot;:11623675,&quot;primary_user_id&quot;:98078198,&quot;theme_var_background_pop&quot;:&quot;#99A2F1&quot;,&quot;created_at&quot;:&quot;2023-12-20T10:59:33.318Z&quot;,&quot;email_from_name&quot;:&quot;Addy and Hassan from Large Scale Apps&quot;,&quot;copyright&quot;:&quot;Addy Osmani and Hassan Djirdeh&quot;,&quot;founding_plan_name&quot;:null,&quot;community_enabled&quot;:true,&quot;invite_only&quot;:false,&quot;payments_state&quot;:&quot;disabled&quot;,&quot;language&quot;:null,&quot;explicit&quot;:false,&quot;homepage_type&quot;:&quot;newspaper&quot;,&quot;is_personal_mode&quot;:false}},{&quot;id&quot;:2224891,&quot;user_id&quot;:11623675,&quot;publication_id&quot;:2209631,&quot;role&quot;:&quot;admin&quot;,&quot;public&quot;:true,&quot;is_primary&quot;:false,&quot;publication&quot;:{&quot;id&quot;:2209631,&quot;name&quot;:&quot;Deep Voice&quot;,&quot;subdomain&quot;:&quot;deepvoice&quot;,&quot;custom_domain&quot;:null,&quot;custom_domain_optional&quot;:false,&quot;hero_text&quot;:&quot;A newsletter on how to get more motivated. Brought to you by Addy Osmani.&quot;,&quot;logo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/328afbab-a375-4ffd-ac83-40300eefc225_1280x1280.png&quot;,&quot;author_id&quot;:11623675,&quot;primary_user_id&quot;:null,&quot;theme_var_background_pop&quot;:&quot;#786CFF&quot;,&quot;created_at&quot;:&quot;2023-12-28T20:05:46.081Z&quot;,&quot;email_from_name&quot;:null,&quot;copyright&quot;:&quot;Addy Osmani&quot;,&quot;founding_plan_name&quot;:&quot;Founding Member&quot;,&quot;community_enabled&quot;:true,&quot;invite_only&quot;:false,&quot;payments_state&quot;:&quot;enabled&quot;,&quot;language&quot;:null,&quot;explicit&quot;:false,&quot;homepage_type&quot;:&quot;newspaper&quot;,&quot;is_personal_mode&quot;:false}}],&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:100,&quot;status&quot;:{&quot;bestsellerTier&quot;:100,&quot;subscriberTier&quot;:null,&quot;leaderboard&quot;:null,&quot;vip&quot;:false,&quot;badge&quot;:{&quot;type&quot;:&quot;bestseller&quot;,&quot;tier&quot;:100},&quot;paidPublicationIds&quot;:[],&quot;subscriber&quot;:null}}],&quot;utm_campaign&quot;:null,&quot;belowTheFold&quot;:true,&quot;type&quot;:&quot;newsletter&quot;,&quot;language&quot;:&quot;en&quot;,&quot;source&quot;:null}" data-component-name="EmbeddedPostToDOM"><a class="embedded-post" native="true" href="https://addyo.substack.com/p/the-70-problem-hard-truths-about?utm_source=substack&amp;utm_campaign=post_embed&amp;utm_medium=web"><div class="embedded-post-header"><img class="embedded-post-publication-logo" src="https://substackcdn.com/image/fetch/$s_!8WxC!,w_56,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3704470-b6d5-48a9-a9d1-564bd833fc5c_1280x1280.png" loading="lazy"><span class="embedded-post-publication-name">Elevate</span></div><div class="embedded-post-title-wrapper"><div class="embedded-post-title">The 70% problem: Hard truths about AI-assisted coding</div></div><div class="embedded-post-body">After spending the last few years embedded in AI-assisted development, I've noticed a fascinating pattern. While engineers report being dramatically more productive with AI, the actual software we use daily doesn&#8217;t seem like it&#8217;s getting noticeably better. What's going on here&#8230;</div><div class="embedded-post-cta-wrapper"><span class="embedded-post-cta">Read more</span></div><div class="embedded-post-meta">2 years ago &#183; 1514 likes &#183; 75 comments &#183; Addy Osmani</div></a></div><p>AI should avoid cleverness. If a simpler solution exists, pick it.</p><h3>&#9989; Don&#8217;t Narrate, Ship</h3><blockquote><p>&#8220;Avoid over-documenting implementation details or rationale unless explicitly required.&#8221;</p></blockquote><p>AI-generated README essays? Save it. We want user-facing docs, not ChatGPT&#8217;s inner monologue.</p><h2>The Bigger Picture</h2><p>Now, every AI task respects my <strong>GitOps-first</strong>, config-driven architecture. New flags go in <code>runner.yaml</code>, not <code>os.getenv</code>. Every schema change comes with a schema update and test. Our platform evolved. Our AI did too.</p><p><code>.cursor/rules</code> isn&#8217;t just instructions &#8211; it&#8217;s <strong>culture</strong>.</p><h2>AI Engineers Need Culture Too</h2><p>We wouldn&#8217;t onboard a junior dev by saying &#8220;just read the codebase.&#8221; We give them docs, a buddy, some rules. We didn&#8217;t do that for our AI&#8230; until it started acting like a junior dev with caffeine poisoning.</p><p><code>.cursor/rules</code> became my AI onboarding playbook. It cut entropy, limited scope creep, and gave the AI the same values we hold. And when your AI shares your values? It becomes a real teammate.</p><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://blog.dativo.io/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Data, Engineering, and Beyond! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Asset-Centric Orchestration: Focus on "What," Not "How"]]></title><description><![CDATA[Why data engineering is stuck in a binary deadlock]]></description><link>https://blog.dativo.io/p/asset-centric-orchestration-focus</link><guid isPermaLink="false">https://blog.dativo.io/p/asset-centric-orchestration-focus</guid><dc:creator><![CDATA[Sergey]]></dc:creator><pubDate>Tue, 30 Dec 2025 14:10:56 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!qTVw!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F887dde6a-fdad-43b8-8b5d-48d12ea61158_914x545.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>I've built data platforms for years, and one thing is clear: modern "no-code" ingestion tools promise the moon but often leave you hanging. The current research establishes that data engineering is stuck in a <strong>binary deadlock</strong>. Organizations must choose between the <strong>Managed Service Tax</strong> (convenience at the cost of opaque billing) and the <strong>Operational Tax</strong> (flexibility at the cost of human capital).</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!qTVw!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F887dde6a-fdad-43b8-8b5d-48d12ea61158_914x545.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!qTVw!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F887dde6a-fdad-43b8-8b5d-48d12ea61158_914x545.jpeg 424w, https://substackcdn.com/image/fetch/$s_!qTVw!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F887dde6a-fdad-43b8-8b5d-48d12ea61158_914x545.jpeg 848w, https://substackcdn.com/image/fetch/$s_!qTVw!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F887dde6a-fdad-43b8-8b5d-48d12ea61158_914x545.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!qTVw!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F887dde6a-fdad-43b8-8b5d-48d12ea61158_914x545.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!qTVw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F887dde6a-fdad-43b8-8b5d-48d12ea61158_914x545.jpeg" width="914" height="545" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/887dde6a-fdad-43b8-8b5d-48d12ea61158_914x545.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:545,&quot;width&quot;:914,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:183676,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://blog.dativo.io/i/182860362?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F887dde6a-fdad-43b8-8b5d-48d12ea61158_914x545.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!qTVw!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F887dde6a-fdad-43b8-8b5d-48d12ea61158_914x545.jpeg 424w, https://substackcdn.com/image/fetch/$s_!qTVw!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F887dde6a-fdad-43b8-8b5d-48d12ea61158_914x545.jpeg 848w, https://substackcdn.com/image/fetch/$s_!qTVw!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F887dde6a-fdad-43b8-8b5d-48d12ea61158_914x545.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!qTVw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F887dde6a-fdad-43b8-8b5d-48d12ea61158_914x545.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">A Portrait of Guan Zhong from a segment of Wu family shrines stone-relief.</figcaption></figure></div><p><em>Two thousand years ago, the Chinese statesman <a href="https://en.wikipedia.org/wiki/Guan_Zhong">Guan Zhong</a> faced a familiar dilemma: tax people directly and risk revolt, or lower taxes and starve the state. His solution was neither. He eliminated most direct taxes entirely and moved revenue into infrastructure&#8212;salt, iron, trade&#8212;making taxation predictable, indirect, and almost invisible.</em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://blog.dativo.io/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Data, Engineering, and Beyond! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>Modern data ingestion is stuck in the same deadlock. Managed services impose a hidden tax through opaque pricing and resync shocks. Code-first pipelines impose an operational tax through human toil. As a engineering leader who has faced late-night firefighting and &#8220;data swamps,&#8221; I&#8217;ve realized that the era of fragmented &#8220;best-of-breed&#8221; tooling is concluding, replaced by integrated platforms that offer streamlined workflows at the cost of increasing vendor lock-in and financial unpredictability.I believe the industry shall follow the third path: shifting cost upstream into structure, specs, and contracts&#8212;so ingestion stops taxing people at all.</p><h2><strong>The Binary Deadlock: Fivetran, dbt, and the &#8220;Fusion&#8221; Shift</strong></h2><p>The October 2025 merger between Fivetran and dbt Labs created a unified big data ETL player. This signals a shift toward integrated stacks where ingestion and transformation are unified. However, this consolidation brings risks for all the data folks:</p><ul><li><p><strong>The Licensing Divergence:</strong> While dbt Core remains Apache 2.0, the "future vision" for transformation lies in <strong>dbt Fusion</strong>. Written in Rust, <a href="https://www.getdbt.com/blog/new-code-new-license-understanding-the-new-license-for-the-dbt-fusion-engine">Fusion is licensed under the </a><strong><a href="https://www.getdbt.com/blog/new-code-new-license-understanding-the-new-license-for-the-dbt-fusion-engine">Elastic License v2 (ELv2)</a></strong>&#8212;a "source-available" license that prohibits using it to provide managed services to others.</p></li><li><p><strong>The Innovation Tax:</strong> As dbt Core enters <a href="https://www.tobikodata.com/blog/dbt-fusion-death-of-dbt-core">"maintenance mode"</a> innovation is reserved for proprietary engines, forcing organizations into a "dangerously dependent" relationship with a single vendor.</p></li></ul><h2><strong>GUI-First vs. Specification-First: The Hidden Operational Costs</strong></h2><p>Airbyte and Fivetran shine with slick UIs. But that magic comes at the cost of control. When engineering talent spends <strong><a href="https://ctomagazine.com/tackling-tech-debt-boosts-agility-data-engineering/">30% to 40% of their time "firefighting"</a></strong>&#8212;fixing broken pipelines and resolving sync failures&#8212;the long-term health of the platform enters a state of terminal decline.</p><ul><li><p><strong>Ingestion Debt:</strong> According to the <a href="https://ctomagazine.com/tackling-tech-debt-boosts-agility-data-engineering/">2024 Global Data Engineering Report</a>, <strong>64% of data leaders</strong> report that technical debt significantly limits their ability to achieve business goals.</p></li><li><p><strong>The &#8220;Fragility Loop&#8221;:</strong> Driven by pressure to deliver, engineers adopt brittle scripts that create a &#8220;data death cycle,&#8221; where every minor change requires manual oversight and eventually <a href="https://datalere.com/articles/two-realities-behind-data-engineering-delays">consumes 60% to 80% of the team&#8217;s capacity</a>.</p></li></ul><h2><strong>&#8220;Open-Source&#8221; Connectors: the limits and  solution</strong></h2><p>Meltano and the Singer ecosystem offer a code-first illusion, but the reality is a maintenance nightmare of outdated taps. To solve the brittleness of  ad-hoc fixing the single opensource provider, I advocate for <strong>Spec-Driven Data pipeline</strong>, which<strong> </strong>declares expected input and desired output. The connectors shall be serving as purely technical providers rather:</p><ul><li><p><strong>Precision over Prompting:</strong> Spec-Driven Development should use formal, machine-readable specifications as a single source of truth for both input and output. This structured collaboration aims for <strong>95% or higher accuracy</strong> in implementing specs on the first attempt.</p></li><li><p><strong>Spec-Driven Workflow:</strong> By moving intellectual effort &#8220;upstream&#8221; to the <strong>Specify</strong> and <strong>Plan</strong> phases, teams capture the &#8220;why&#8221; behind technical choices, preventing &#8220;intent-vs-implementation drift&#8221;. Of course the tehnical aspects of connectors and data movers are still present, but they are not so important and can be replaced with other providers if there is such need. </p></li></ul><h2><strong>Data Contracts: Beyond &#8220;Bring First-Decide Later&#8221;</strong></h2><p>The traditional approach of ingesting raw data and transforming it later leads to data swamps. Data quality issues cost organizations an average of <strong><a href="https://angrynerds.co/blog/data-engineering-solutions-4-business-challenges/">$12.9 million annually</a></strong><a href="https://angrynerds.co/blog/data-engineering-solutions-4-business-challenges/">.</a></p><ul><li><p><strong>Fail-Closed Validation:</strong> We declare machine-readable YAML contracts at the source. If a source system change violates the contract, the deployment is rejected&#8212;a <strong>&#8220;fail-closed&#8221;</strong> gate that prevents downstream pollution.</p></li><li><p><strong>Schema-as-Code:</strong> By versioning every asset's schema in Git and running strict validation at ingestion time, we would treat data movement with the same rigor as application code.</p></li></ul><h2><strong>Asset-Centric Orchestration: Focus on &#8220;What,&#8221; Not &#8220;How&#8221;</strong></h2><p>Traditional orchestrators like Airflow are task-centric, focusing on workflow execution. We should utilize an <strong>asset-centric</strong> approach in order to keep an eye on the thing which is really important - &#8220;What&#8221; we are expecting to get as the result of data pipeline.</p><ul><li><p><strong>Native Lineage:</strong> In an asset-centric model, the focus have be on the data products produced. Then lineage is captured automatically in a unified graph, making it easier to trace origin and transformation.</p></li><li><p><strong>One Job per one Asset:</strong> If we would enforce a one-asset-per-job design pattern, this will simplify retry logic and ensures that if one of fifty tables fails, we only re-run that specific job rather than retrying a massive, entangled pipeline.</p></li></ul><p></p><h2><strong>The &#8220;Third Way&#8221;</strong></h2><p>After slogging through these limitations, I built <a href="https://github.com/dativo-io/dativo-ingest">dativo-ingest</a> to test the hypothesis of breaking the binary deadlock.</p><ul><li><p><strong>Headless &amp; Config-Driven:</strong> No UI needed; pipelines are defined in YAML under GitOps, making deployments auditable.</p></li><li><p><strong>Lakehouse-Native:</strong> It writes Apache Iceberg tables directly and update metadata via Nessie, ensuring ACID semantics and propagation of FinOps tags directly into table properties.</p></li><li><p><strong>Metadata-Driven Production Readiness:</strong> I specifically designed it for high-scale environments like Databricks Lakeflow, generating production-ready code automatically from formal specs.</p></li></ul><p>I am going to continue posting about my findings, but one more time to underline - Dativo Ingest isn&#8217;t just about connectors - they are expremely replacable in the YAML configurations. The ineventables there are constraints&#8212;like upfront schemas and tags&#8212;because they force the discipline required to escape the data death cycle and reclaim technical agency.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://blog.dativo.io/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Data, Engineering, and Beyond! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item></channel></rss>