<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Blog | Ben Benhemo</title><link>https://benbenhemo.com/post/</link><atom:link href="https://benbenhemo.com/post/index.xml" rel="self" type="application/rss+xml"/><description>Blog</description><generator>Hugo Blox Builder (https://hugoblox.com)</generator><language>en-us</language><image><url>https://benbenhemo.com/media/icon_hu9e1d2b86e2bb2877819b4fa069da1ee7_107810_512x512_fill_lanczos_center_3.png</url><title>Blog</title><link>https://benbenhemo.com/post/</link></image><item><title>The Security Archipelago: Why Your Tools are Sovereign States in a Lawless Sea</title><link>https://benbenhemo.com/post/the_security_archipelago/</link><pubDate>Tue, 27 Jan 2026 00:00:00 +0000</pubDate><guid>https://benbenhemo.com/post/the_security_archipelago/</guid><description>&lt;h3 id="i--the-specialists-blindness">&lt;strong>I — The Specialist’s Blindness&lt;/strong>&lt;/h3>
&lt;p>If you’re anything like me, you were taught that security is a game of coverage.&lt;/p>
&lt;p>We were told that if we just have enough domain expert tools, an EDR for the endpoint, a CSPM for the cloud, a SIEM for monitoring, the &amp;ldquo;gaps&amp;rdquo; would disappear. We built a stack of &amp;ldquo;Best in Class&amp;rdquo; sovereigns.&lt;/p>
&lt;p>But there is a paradox at the heart of the modern security stack: The more specialized our tools become, the more fragmented our reality feels.&lt;/p>
&lt;p>We didn&amp;rsquo;t build a unified defense. We built an &lt;strong>Archipelago of Data.&lt;/strong>&lt;/p>
&lt;h3 id="ii--you-arent-secure-because-you-have-coverage-youre-blind-because-you-have-silos">&lt;strong>II — You Aren&amp;rsquo;t Secure Because You Have Coverage, You&amp;rsquo;re Blind Because You Have Silos&lt;/strong>&lt;/h3>
&lt;p>Think of your security stack as a chain of isolated islands. Each island is a sovereign state with its own laws, its own language, and its own guarded borders.&lt;/p>
&lt;ul>
&lt;li>&lt;strong>The Island of Endpoint (EDR)&lt;/strong> speaks the language of processes, kernel events, and file hashes.&lt;/li>
&lt;li>&lt;strong>The Island of Identity (IdP)&lt;/strong> speaks the language of user profiles, MFA status, and group memberships.&lt;/li>
&lt;li>&lt;strong>The Island of Infrastructure (Cloud)&lt;/strong> speaks the language of resource exposure, vulnerabilities, and attack paths.&lt;/li>
&lt;/ul>
&lt;p>Each island is a &amp;ldquo;Domain Expert&amp;rdquo;. Within its own borders each tool is nearly perfect. But the business doesn&amp;rsquo;t happen on the islands, it happens in the transit between them.&lt;/p>
&lt;p>Because these islands refuse to speak the same language the burden of correlation falls on the human analyst. When a potential issue is flagged your team is forced to act as a &lt;strong>manual ferry&lt;/strong>.&lt;/p>
&lt;p>For example when you ask a cross platform question:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" class="chroma">&lt;code class="language-fallback" data-lang="fallback">&lt;span class="line">&lt;span class="cl">&amp;#34;Which user is an admin in AWS despite using unmanaged devices,
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">has no MFA configuration in the IdP, and is being targeted
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">by a high volume of malicious emails in the last month?&amp;#34;
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>The archipelago falls silent. No single island has the answer because no single island owns the &amp;ldquo;connective tissue&amp;rdquo; of the data.&lt;/p>
&lt;p>You are paying your most expensive talent to do the work of correlation. This isn&amp;rsquo;t security analysis. It&amp;rsquo;s a &lt;strong>Translation Tax&lt;/strong>.&lt;/p>
&lt;h3 id="iii--the-outcome-achieving-architectural-visibility">&lt;strong>III — The Outcome: Achieving Architectural Visibility&lt;/strong>&lt;/h3>
&lt;p>The goal of bridging the archipelago isn&amp;rsquo;t just to catch risks faster, it&amp;rsquo;s to fundamentally change the way your team works. When you move from &lt;strong>Linear Observation&lt;/strong> (looking at one tool at a time) to &lt;strong>Architectural Visibility&lt;/strong> (seeing the whole map), the whole thing changes.&lt;/p>
&lt;p>The value of extracting and linking these data points can help you to gain the ability to:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Map the &amp;ldquo;Blast Radius&amp;rdquo;:&lt;/strong> Instantly see what an identity can touch across the entire environment before it is ever compromised.&lt;/li>
&lt;li>&lt;strong>Prioritize the Human Element:&lt;/strong> Focus on the users who are more likely to be targeted by external threats, rather than treating every identity with equal weight.&lt;/li>
&lt;li>&lt;strong>Eliminate the Tax:&lt;/strong> Reclaim the hours lost to manual correlation and redirect them toward high value security visibility.&lt;/li>
&lt;/ul>
&lt;h3 id="iv--conclusion-build-the-bridge">&lt;strong>IV — Conclusion: Build the Bridge&lt;/strong>&lt;/h3>
&lt;p>The life of a security teams shouldn&amp;rsquo;t be spent in the gaps between platforms. The future belongs to those who recognize that domain expertise is only half the battle. The other half is the connectivity that turns experts into a system.&lt;/p>
&lt;p>You don&amp;rsquo;t have to get more islands. Just start building the bridges.&lt;/p></description></item><item><title>Lightweight Reachability System: GitLab Knowledge Graph + AI Agents</title><link>https://benbenhemo.com/post/lightweight-reachability/</link><pubDate>Sat, 04 Oct 2025 00:00:00 +0000</pubDate><guid>https://benbenhemo.com/post/lightweight-reachability/</guid><description>&lt;h2 id="the-problem-not-every-sca-vulnerability-is-exploitable">The Problem: Not Every SCA Vulnerability is Exploitable&lt;/h2>
&lt;p>When &lt;a href="https://nvd.nist.gov/vuln/detail/CVE-2025-29927" target="_blank" rel="noopener">CVE-2025-29927&lt;/a> dropped for the Next.js library, many appsec teams worldwide immediately started working on a fix.
Dependency scanners generated a vast amount of security findings across companies codebases that had the vulnerable version of the package.&lt;/p>
&lt;p>The reason many false positive alerts were generated was due to a technical aspect that most traditional dependency scanners miss: &lt;strong>Vulnerability != Exploitability&lt;/strong>.&lt;/p>
&lt;p>The Next.js CVE perfectly illustrates this gap. From an exploitability perspective this vulnerability is only relevant where:&lt;/p>
&lt;ul>
&lt;li>Self-hosted deployments using &lt;code>next start&lt;/code> with &lt;code>output: standalone&lt;/code>&lt;/li>
&lt;li>Middleware file &lt;code>(middleware.js/middleware.ts)&lt;/code> exists and is actively used&lt;/li>
&lt;li>Middleware performs critical security operations like authentication or authorization checks&lt;/li>
&lt;/ul>
&lt;p>Yet every traditional dependency scanner will flag any application using affected Next.js versions, regardless of deployment context or middleware usage. This creates a massive false positive problem that burns security &amp;amp; engineering hours and creates alert fatigue.&lt;/p>
&lt;h2 id="why-reachability-analysis-matters">Why Reachability Analysis Matters&lt;/h2>
&lt;p>Traditional SCA (Software Composition Analysis) tools operate on a simple binary: They extract the SBOM ➡️ Compare versions to a vulnerability database ➡️ Flag any match as risk.&lt;/p>
&lt;p>This simplistic approach results in a high rate of false positives, as it doesn&amp;rsquo;t consider whether the vulnerable code is actually reachable or exploitable in your specific context.
Because real exploitation requires:&lt;/p>
&lt;ol>
&lt;li>&lt;strong>Vulnerable code must be imported&lt;/strong> - Is the vulnerable function/module actually imported?&lt;/li>
&lt;li>&lt;strong>Reachable execution path&lt;/strong> - Does the call graph show that application flow can reach the vulnerable code?&lt;/li>
&lt;li>&lt;strong>Client controlled input&lt;/strong> - Can external input influence the vulnerable code path?&lt;/li>
&lt;/ol>
&lt;h2 id="gitlabs-knowledge-graph-tool-overview">GitLab&amp;rsquo;s Knowledge Graph: Tool Overview&lt;/h2>
&lt;p>GitLab recently released a &lt;a href="https://gitlab-org.gitlab.io/rust/knowledge-graph/" target="_blank" rel="noopener">Knowledge Graph tool&lt;/a> that transforms your codebase into a queryable graph database. It understands:&lt;/p>
&lt;ul>
&lt;li>Function call hierarchies&lt;/li>
&lt;li>Import dependencies&lt;/li>
&lt;li>Variable flow&lt;/li>
&lt;li>Cross-file relationships&lt;/li>
&lt;/ul>
&lt;img src="knowledge-graph.jpg" alt="GitLab Knowledge Graph" style="width: 100%; max-width: 100%; height: auto;" />
&lt;p>The special thing about this tool is that it can be easily integrated into your AI agents &amp;amp; LLMs by providing a local MCP setup, which helps you query your codebase for code insights.&lt;/p>
&lt;div class="flex px-4 py-3 rounded-md bg-primary-100 dark:bg-primary-900">
&lt;span class="pr-3 pt-1 text-primary-400">
&lt;svg height="24" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24">&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="m11.25 11.25l.041-.02a.75.75 0 0 1 1.063.852l-.708 2.836a.75.75 0 0 0 1.063.853l.041-.021M21 12a9 9 0 1 1-18 0a9 9 0 0 1 18 0m-9-3.75h.008v.008H12z"/>&lt;/svg>
&lt;/span>
&lt;span class="dark:text-neutral-300">While this tool is primarily designed to assist engineers during development, I realized its capabilities could be redirected to help AppSec engineers investigate SCA vulnerabilities.&lt;/span>
&lt;/div>
&lt;div style="position: relative; padding-bottom: 56.25%; height: 0; overflow: hidden;">
&lt;iframe src="https://www.youtube.com/embed/wL6-m5_2FH8" style="position: absolute; top: 0; left: 0; width: 100%; height: 100%; border:0;" allowfullscreen title="YouTube Video">&lt;/iframe>
&lt;/div>
&lt;h2 id="quick-example-cve-2024-47081-analysis">Quick Example: CVE-2024-47081 Analysis&lt;/h2>
&lt;p>To demonstrate this approach I analyzed a sample repository for SCA findings. From the initial traditional scan I identified the requests library at version 2.32.3 which has CVE-2024-47081: a credential leakage vulnerability in Python&amp;rsquo;s requests library.
This vulnerability allows .netrc credential leakage when processing malicious URLs, but our analysis revealed:&lt;/p>
&lt;ul>
&lt;li>3 files using the requests library: github_collector.py, npm_collector.py, and smithery_collector.py&lt;/li>
&lt;li>All requests calls use hardcoded URL templates with validated string formatting&lt;/li>
&lt;li>No arbitrary URL injection vectors found in the codebase&lt;/li>
&lt;/ul>
&lt;p>This analysis took &lt;strong>&amp;lt; 30 seconds&lt;/strong> using GitLab&amp;rsquo;s Knowledge Graph + Claude, compared to a long time of manual code review. The system correctly identified that while the vulnerable package was present, the specific conditions for exploitation were not met.&lt;/p>
&lt;img src="screenshot-requests.jpg" alt="CVE Analysis Screenshot" style="width: 100%; max-width: 100%; height: auto;" />
&lt;h2 id="automating-the-process">Automating the Process&lt;/h2>
&lt;p>Taking this a step further, we can essentially automate the process to help with remediations by creating a job that takes the SCA findings from our traditional scanners and uses the knowledge graph along with an AI agent to review them. It will automatically generate a reachability report for each repository.&lt;/p>
&lt;h3 id="architecture-overview">Architecture Overview&lt;/h3>
&lt;div class="mermaid">graph LR
A[SCA Scanner Findings] --> B[LLM + Knowledge Graph]
B --> C[Call Graph Analysis]
C --> D[Reachability Report]
E[Your Codebase] --> B
&lt;/div>
&lt;h3 id="example-the-reachability-analysis-prompt">Example: The Reachability Analysis Prompt&lt;/h3>
&lt;p>Here&amp;rsquo;s the prompt template that powers the analysis:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" class="chroma">&lt;code class="language-markdown" data-lang="markdown">&lt;span class="line">&lt;span class="cl">You are a security researcher analyzing CVE reachability.
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">Given CVE: [CVE-ID]
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">Vulnerable component: [package@version]
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">Vulnerability details: [description]
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">Target Project: [project]
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">Using the GitLab Knowledge Graph, determine:
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="k">1.&lt;/span> Is the vulnerable package imported anywhere?
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="k">2.&lt;/span> What are the call paths to vulnerable functions?
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="k">3.&lt;/span> Are these paths reachable from external entry points?
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="k">4.&lt;/span> What conditions must be met for exploitation?
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">Provide a reachability verdict: REACHABLE | UNREACHABLE
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;h2 id="beyond-cves-advanced-security-applications">Beyond CVEs: Advanced Security Applications&lt;/h2>
&lt;p>The knowledge graph isn&amp;rsquo;t just for investigating CVEs, it opens up entirely new possibilities for answering complex security questions that traditional SAST and SCA tools can&amp;rsquo;t address. For example:&lt;/p>
&lt;h3 id="finding-authentication-bypasses">Finding Authentication Bypasses&lt;/h3>
&lt;div class="highlight">&lt;pre tabindex="0" class="chroma">&lt;code class="language-fallback" data-lang="fallback">&lt;span class="line">&lt;span class="cl">&amp;#34;Show me all code flows that reach database queries without passing through an authentication check&amp;#34;
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;h3 id="tracking-pii-flow">Tracking PII Flow&lt;/h3>
&lt;div class="highlight">&lt;pre tabindex="0" class="chroma">&lt;code class="language-fallback" data-lang="fallback">&lt;span class="line">&lt;span class="cl">&amp;#34;Trace all paths where PII data flows from API input to external services&amp;#34;
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;h3 id="detecting-ssrf-patterns">Detecting SSRF Patterns&lt;/h3>
&lt;div class="highlight">&lt;pre tabindex="0" class="chroma">&lt;code class="language-fallback" data-lang="fallback">&lt;span class="line">&lt;span class="cl">&amp;#34;Find all places where user input can influence URL parameters in HTTP requests&amp;#34;
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>&lt;strong>The combination of GitLab&amp;rsquo;s Knowledge Graph and AI agents enables us to perform deep level analysis of SCA vulnerabilities in a budget friendly manner, using this open source tool combined with our AI agents. This approach not only saves time and reduces alert fatigue, but it also opens up new possibilities for automated security analysis without the need for expensive commercial tools.&lt;/strong>&lt;/p>
&lt;hr>
&lt;small>
&lt;p>&lt;strong>References:&lt;/strong>&lt;/p>
&lt;ol>
&lt;li>CVE-2025-29927 - Next.js Authorization Bypass Vulnerability | &lt;a href="https://nvd.nist.gov/vuln/detail/CVE-2025-29927" target="_blank" rel="noopener">National Vulnerability Database&lt;/a>&lt;/li>
&lt;li>CVE-2025-29927 Deep Dive | &lt;a href="https://jfrog.com/blog/cve-2025-29927-next-js-authorization-bypass/" target="_blank" rel="noopener">JFrog Security Research&lt;/a>&lt;/li>
&lt;li>GitLab Knowledge Graph MCP Server | &lt;a href="https://gitlab-org.gitlab.io/rust/knowledge-graph/" target="_blank" rel="noopener">Official Documentation&lt;/a>&lt;/li>
&lt;li>CVE-2024-47081 - Python Requests Credential Leakage | &lt;a href="https://github.com/advisories/GHSA-9wx4-h78v-vm56" target="_blank" rel="noopener">GitHub Security Advisory&lt;/a>&lt;/li>
&lt;/ol>
&lt;/small></description></item><item><title>The Three Pillars For Container Security</title><link>https://benbenhemo.com/post/container_security_three_pillars/</link><pubDate>Sun, 18 Aug 2024 00:00:00 +0000</pubDate><guid>https://benbenhemo.com/post/container_security_three_pillars/</guid><description>&lt;p>When you are surfing the web, you can find countless approaches on how to secure containers. There are plenty best practices available, but this abundance of information can often be misleading and confusing.&lt;/p>
&lt;p>This dedicated guide aims to cut through the noise and present a focused perspective on what I believe are the three most critical pillars of container security. These pillars are essential for anyone in the security field to understand and implement when securing containers. This guide assumes that you already have a basic understanding of what Docker and containers are.&lt;/p>
&lt;p>So, let’s get started.&lt;/p>
&lt;h2 id="tldr-of-the-three-pillars">&lt;strong>TL;DR of the Three Pillars&lt;/strong>&lt;/h2>
&lt;div class="flex px-4 py-3 rounded-md bg-primary-100 dark:bg-primary-900">
&lt;span class="pr-3 pt-1 text-primary-400">
&lt;svg height="24" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24">&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="m11.25 11.25l.041-.02a.75.75 0 0 1 1.063.852l-.708 2.836a.75.75 0 0 0 1.063.853l.041-.021M21 12a9 9 0 1 1-18 0a9 9 0 0 1 18 0m-9-3.75h.008v.008H12z"/>&lt;/svg>
&lt;/span>
&lt;span class="dark:text-neutral-300">💡 &lt;strong>Securing Docker Images, Securing Kubernetes, Runtime Security.&lt;/strong>&lt;/span>
&lt;/div>
&lt;br>
&lt;h2 id="1securing-docker-images">1️⃣ Securing Docker Images&lt;/h2>
&lt;p>Securing your Dockerfile is the first process you should follow.&lt;/p>
&lt;h3 id="some-basics-understanding-the-docker-build-process">&lt;strong>Some Basics: Understanding the Docker Build Process&lt;/strong>&lt;/h3>
&lt;p>A Docker build process involves creating a Docker image from a Dockerfile. A Dockerfile is essentially a blueprint containing instructions that define the environment and the steps required to build the application.&lt;/p>
&lt;p>During the build process, Docker reads these instructions, executes them, and packages the results into an image. This image can then be run as a container, ensuring a consistent and reproducible environment for applications.&lt;/p>
&lt;p>
&lt;figure id="figure-docker">
&lt;div class="flex justify-center ">
&lt;div class="w-100" >&lt;img alt="Docker" srcset="
/post/container_security_three_pillars/docker_1_hu80fd3cb0129969eda225828bcfb815e1_174945_5fec513bfed7b202d916e1b1c4ff7a01.webp 400w,
/post/container_security_three_pillars/docker_1_hu80fd3cb0129969eda225828bcfb815e1_174945_cbdebf1033f00e6365c50fd75b68b12f.webp 760w,
/post/container_security_three_pillars/docker_1_hu80fd3cb0129969eda225828bcfb815e1_174945_1200x1200_fit_q95_h2_lanczos.webp 1200w"
src="https://benbenhemo.com/post/container_security_three_pillars/docker_1_hu80fd3cb0129969eda225828bcfb815e1_174945_5fec513bfed7b202d916e1b1c4ff7a01.webp"
width="760"
height="305"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;figcaption>
Docker
&lt;/figcaption>&lt;/figure>
&lt;/p>
&lt;div class="flex px-4 py-3 rounded-md bg-primary-100 dark:bg-primary-900">
&lt;span class="pr-3 pt-1 text-primary-400">
&lt;svg height="24" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24">&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="m11.25 11.25l.041-.02a.75.75 0 0 1 1.063.852l-.708 2.836a.75.75 0 0 0 1.063.853l.041-.021M21 12a9 9 0 1 1-18 0a9 9 0 0 1 18 0m-9-3.75h.008v.008H12z"/>&lt;/svg>
&lt;/span>
&lt;span class="dark:text-neutral-300">💡 A great explanation of the build process can be found here: &lt;a href="https://iximiuz.com/en/posts/you-need-containers-to-build-an-image/" target="_blank" rel="noopener">https://iximiuz.com/en/posts/you-need-containers-to-build-an-image/&lt;/a>&lt;/span>
&lt;/div>
&lt;h3 id="dockerfile-security">Dockerfile Security&lt;/h3>
&lt;p>From what I have seen so far, the three most important aspects of securing Dockerfiles are:&lt;/p>
&lt;ul>
&lt;li>Pulling a secure base image.&lt;/li>
&lt;li>Securing additional dependencies and packages.&lt;/li>
&lt;li>Ensuring containers do not run as root.&lt;/li>
&lt;/ul>
&lt;h4 id="base-image">&lt;strong>Base Image&lt;/strong>&lt;/h4>
&lt;p>The first layer in a Dockerfile is the &lt;strong>base image&lt;/strong>. This is the foundation upon which all other layers are built. The base image typically comes from an external registry.&lt;/p>
&lt;p>&lt;strong>It is important to check the current vulnerability state of an image in the external registry. For instance, Docker Hub provides information on the existing vulnerabilities in an image.&lt;/strong>&lt;/p>
&lt;ul>
&lt;li>
&lt;p>In the following example, you can see an node:10.0.0 image sourced from Docker Hub. The image shows a significant number of vulnerabilities and was last pushed six years ago.&lt;/p>
&lt;/li>
&lt;li>
&lt;p>This indicates that the image is outdated and vulnerable. Therefore, it is advisable to upgrade to a newer version or change the image to a different tag to ensure better security.&lt;/p>
&lt;p>
&lt;figure id="figure-node">
&lt;div class="flex justify-center ">
&lt;div class="w-100" >&lt;img alt="Node" srcset="
/post/container_security_three_pillars/node_hu409f8433c03571b3276997f46f698d2f_133438_8dcb88a852f7aa8143e612da035a83a0.webp 400w,
/post/container_security_three_pillars/node_hu409f8433c03571b3276997f46f698d2f_133438_4025470cea023796f3447baa510a25b6.webp 760w,
/post/container_security_three_pillars/node_hu409f8433c03571b3276997f46f698d2f_133438_1200x1200_fit_q95_h2_lanczos.webp 1200w"
src="https://benbenhemo.com/post/container_security_three_pillars/node_hu409f8433c03571b3276997f46f698d2f_133438_8dcb88a852f7aa8143e612da035a83a0.webp"
width="760"
height="188"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;figcaption>
Node
&lt;/figcaption>&lt;/figure>
&lt;/p>
&lt;/li>
&lt;/ul>
&lt;h4 id="securing-additional-dependencies-and-packages">Securing additional dependencies and packages.&lt;/h4>
&lt;p>Since a Dockerfile is built using layers and starting with a base image, developers will often need to add dependencies and packages on top of that. These dependencies and packages are crucial for the functionality of the application.&lt;/p>
&lt;p>It is essential to ensure that all dependencies and packages are up-to-date and secure.&lt;/p>
&lt;div class="flex px-4 py-3 rounded-md bg-primary-100 dark:bg-primary-900">
&lt;span class="pr-3 pt-1 text-primary-400">
&lt;svg height="24" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24">&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="m11.25 11.25l.041-.02a.75.75 0 0 1 1.063.852l-.708 2.836a.75.75 0 0 0 1.063.853l.041-.021M21 12a9 9 0 1 1-18 0a9 9 0 0 1 18 0m-9-3.75h.008v.008H12z"/>&lt;/svg>
&lt;/span>
&lt;span class="dark:text-neutral-300">💡 For more information on securing dependencies and packages, feel free to read my previous blog: &lt;a href="https://benbenhemo.com/post/securing-dependencies/" target="_blank" rel="noopener">https://benbenhemo.com/post/securing-dependencies/&lt;/a>&lt;/span>
&lt;/div>
&lt;h4 id="ensuring-containers-do-not-run-as-root">Ensuring containers do not run as root&lt;/h4>
&lt;p>&lt;strong>By default, containers run as root&lt;/strong>. Since containers might share the same host, a container running as root will have access to the root of the host node. This means it can potentially access other containers on the same node, posing significant security risks.&lt;/p>
&lt;p>To mitigate these risks, ensure that your Dockerfile is configured to run containers as a non-root user.&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" class="chroma">&lt;code class="language-docker" data-lang="docker">&lt;span class="line">&lt;span class="cl">&lt;span class="c"># Use an official base image&lt;/span>&lt;span class="err">
&lt;/span>&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="err">&lt;/span>&lt;span class="k">FROM&lt;/span>&lt;span class="s"> python:3.9-slim&lt;/span>&lt;span class="err">
&lt;/span>&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="err">
&lt;/span>&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="err">&lt;/span>&lt;span class="c"># Create a non-root user and set permissions&lt;/span>&lt;span class="err">
&lt;/span>&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="err">&lt;/span>&lt;span class="k">RUN&lt;/span> useradd -m myuser&lt;span class="err">
&lt;/span>&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="err">
&lt;/span>&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="err">&lt;/span>&lt;span class="c"># Set the user to the newly created non-root user&lt;/span>&lt;span class="err">
&lt;/span>&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="err">&lt;/span>&lt;span class="k">USER&lt;/span>&lt;span class="s"> myuser&lt;/span>&lt;span class="err">
&lt;/span>&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;h3 id="container-scanning-in-cicd">&lt;strong>Container Scanning in CI/CD&lt;/strong>&lt;/h3>
&lt;p>Implementing and configuring security scanners as part of the CI/CD pipeline is a topic for another day, but since it is a major component of securing Docker images, it should be mentioned in this blog.&lt;/p>
&lt;p>Container scanning in CI/CD pipelines ensures that any vulnerabilities in your Docker images are detected early in the development cycle. This approach helps maintain the security and integrity of your applications by detecting and addressing potential vulnerabilities before they reach production.&lt;/p>
&lt;p>The Trivy documentation is a great resource and can help you gain a more in-depth understanding of the concept: &lt;a href="https://aquasecurity.github.io/trivy/v0.53/" target="_blank" rel="noopener">Trivy Documentation&lt;/a>&lt;/p>
&lt;h2 id="2securing-kubernetes">2️⃣ &lt;strong>Securing Kubernetes&lt;/strong>&lt;/h2>
&lt;p>once your Dockerfile is secure, the next important layer is securing your Kubernetes cluster.&lt;/p>
&lt;p>Kubernetes, an open-source project, is an important tool for deploying, scaling, and managing containerized applications. Its flexibility and power make it a key component in modern infrastructure, but this also means that securing Kubernetes effectively is essential to protect your applications and data from potential threats.&lt;/p>
&lt;ul>
&lt;li>&lt;a href="https://kubernetes.io/docs/home/" target="_blank" rel="noopener">https://kubernetes.io/docs/home/&lt;/a>&lt;/li>
&lt;/ul>
&lt;p>
&lt;figure >
&lt;div class="flex justify-center ">
&lt;div class="w-100" >&lt;img src="https://anthonyspiteri.net/wp-content/uploads/2019/07/k8severywhere.jpg" alt="https://anthonyspiteri.net/wp-content/uploads/2019/07/k8severywhere.jpg" loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;h3 id="securing-the-control-plane">&lt;strong>Securing The Control Plane&lt;/strong>&lt;/h3>
&lt;p>The Kubernetes API server is the heart of your cluster’s control plane, managing all interactions within the cluster. Given its critical role, securing the API server is extremely important. Here are the key steps to ensure its security:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Implement Role-Based Access Control (RBAC):&lt;/strong> Role-Based Access Control (RBAC) is essential for limiting access to the Kubernetes API based on the principle of least privilege. RBAC allows you to define precise roles and permissions for users and applications, ensuring that only authorized entities can interact with the API.&lt;/li>
&lt;li>&lt;strong>Secure API Communication with TLS:&lt;/strong> Ensure that the API server is only accessible over a secure, encrypted connection using TLS. TLS encryption protects the data transmitted between the API server and its clients, preventing unauthorized access and tampering.&lt;/li>
&lt;li>&lt;strong>Regularly Auditing API Server Logs:&lt;/strong> Monitor and audit API server logs to detect unusual or unauthorized activities. Implement logging and monitoring tools to provide visibility into API usage and potential security incidents.&lt;/li>
&lt;/ul>
&lt;p>You can also check: &lt;a href="https://kubernetes.io/docs/concepts/security/controlling-access/" target="_blank" rel="noopener">https://kubernetes.io/docs/concepts/security/controlling-access/&lt;/a>&lt;/p>
&lt;h3 id="pod-security">&lt;strong>Pod Security&lt;/strong>&lt;/h3>
&lt;p>Pods are groups of containers running on the same host, sharing a network stack and other resources. It’s essential to configure your pods to limit their permissions and capabilities, minimizing potential attack surfaces and preventing lateral movement. This can be achieved by:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Setting Pod Security Admission:&lt;/strong> Define a set of conditions that a pod must meet to be accepted into the system. These conditions control aspects such as privilege escalation, host network access, and volume usage.
&lt;ul>
&lt;li>More details can be found here: &lt;a href="https://kubernetes.io/docs/concepts/security/pod-security-admission/" target="_blank" rel="noopener">https://kubernetes.io/docs/concepts/security/pod-security-admission/&lt;/a>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>&lt;strong>Using Security Contexts:&lt;/strong> Security contexts allow you to define privileges and access control settings for a pod or container. For example, you can prevent containers from running as root by setting the runAsNonRoot option. &lt;a href="https://kubernetes.io/docs/concepts/security/pod-security-standards/" target="_blank" rel="noopener">https://kubernetes.io/docs/concepts/security/pod-security-standards/&lt;/a>&lt;/li>
&lt;/ul>
&lt;h3 id="host-security">&lt;strong>Host Security&lt;/strong>&lt;/h3>
&lt;p>Host security is a critical yet often overlooked aspect of securing Kubernetes. The host machines, whether they are bare-metal or virtual machines (VMs), must be properly secured to protect the entire Kubernetes environment.&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Restrict Access:&lt;/strong> Limit access to your host machines by using dedicated VPNs and strict access controls. This helps prevent unauthorized users from accessing the underlying infrastructure.&lt;/li>
&lt;li>&lt;strong>Minimize Attack Surface:&lt;/strong> Only install the necessary components required for running Kubernetes, such as the Kubernetes code, its dependencies (like Docker), and essential supporting features (logging or security tools). This is often referred to as running a “thin OS.” Container-specific distributions like CoreOS Container Linux, RancherOS, or Red Hat Atomic can help minimize the installed code on your hosts and include features like a read-only root filesystem.&lt;/li>
&lt;li>&lt;strong>Use Hardened OS Distributions:&lt;/strong> Consider using container-specific OS distributions or general-purpose Linux distributions that are stripped of unnecessary libraries and tools. This reduces potential vulnerabilities by limiting the software running on the host.&lt;/li>
&lt;/ul>
&lt;div class="flex px-4 py-3 rounded-md bg-primary-100 dark:bg-primary-900">
&lt;span class="pr-3 pt-1 text-primary-400">
&lt;svg height="24" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24">&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="m11.25 11.25l.041-.02a.75.75 0 0 1 1.063.852l-.708 2.836a.75.75 0 0 0 1.063.853l.041-.021M21 12a9 9 0 1 1-18 0a9 9 0 0 1 18 0m-9-3.75h.008v.008H12z"/>&lt;/svg>
&lt;/span>
&lt;span class="dark:text-neutral-300">💡 Discussing Kubernetes security can be an extensive and varied topic. For a comprehensive and informative guide, I recommend reading the book &lt;em>Kubernetes Security&lt;/em> by Liz Rice. It’s an excellent resource that covers a wide range of important topics in depth.&lt;/span>
&lt;/div>
&lt;h2 id="3runtime-security">3️⃣ &lt;strong>Runtime Security&lt;/strong>&lt;/h2>
&lt;p>If you work in the security industry, you’ve probably heard the phrase “Runtime Security.” Although runtime security is not entirely new, it has gained significant adoption in recent years, with companies like Wiz, Upwind Security, and Sweet Security that utilizing this technology. But what exactly is runtime security, and how does it work?&lt;/p>
&lt;h3 id="understanding-runtime-security">&lt;strong>Understanding Runtime Security&lt;/strong>&lt;/h3>
&lt;p>Runtime security refers to the protection of applications while they are running, in that simple but that complex too. It involves monitoring and securing the application in real time to detect and mitigate threats as they occur. This is crucial because even if your application is secure at the build and deploy stages, vulnerabilities can still be exploited at runtime.&lt;/p>
&lt;h4 id="ebpf-the-engine-behind-runtime-security">&lt;strong>eBPF: The Engine Behind Runtime Security&lt;/strong>&lt;/h4>
&lt;p>eBPF (short for Extended Berkeley Packet Filter) is a powerful Linux kernel technology that enables efficient and dynamic tracing of various system events such as network packets, function calls, and system events. Originally, BPF (Berkeley Packet Filter) was developed in the early 1990s to filter network packets in a performant way, primarily in network monitoring tools like tcpdump.&lt;/p>
&lt;h4 id="kernel-vs-user-space-and-how-ebpf-leverages-it">&lt;strong>Kernel vs. User Space and How eBPF Leverages It&lt;/strong>&lt;/h4>
&lt;p>The Linux kernel is the core software layer that sits between applications and the hardware they run on. Applications operate in a restricted environment called &lt;strong>user space&lt;/strong>, where they can’t directly access hardware. Instead, they make system calls to the kernel, which handles these requests.&lt;/p>
&lt;p>eBPF operates in the &lt;strong>kernel space&lt;/strong>, allowing it to monitor and interact with system events in real time. What makes eBPF powerful is that it can do this without modifying the kernel itself, maintaining system stability and security. This capability enables eBPF to provide deep visibility and control, making it a vital tool for real-time security monitoring.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="flex justify-center ">
&lt;div class="w-100" >&lt;img src="https://www.oreilly.com/api/v2/epubs/9781098135119/files/assets/lebp_0101.png" alt="https://www.oreilly.com/api/v2/epubs/9781098135119/files/assets/lebp_0101.png" loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>eBPF technology allows companies and products to monitor and extract crucial insights and data in real time during application runtime. Since every network call and system call originates from the kernel, eBPF operates close to the kernel level, providing deep visibility into system behavior without affecting the actual kernel operations.&lt;/p>
&lt;ul>
&lt;li>A great book about eBPF: &lt;a href="https://isovalent.com/books/learning-ebpf/" target="_blank" rel="noopener">https://isovalent.com/books/learning-ebpf/&lt;/a>&lt;/li>
&lt;li>I believe the best way to learn about runtime security is by getting hands-on experience with it. I highly recommend exploring the Falco project, an open-source security tool that leverages runtime security concepts. &lt;a href="https://falco.org/" target="_blank" rel="noopener">https://falco.org/&lt;/a>&lt;/li>
&lt;/ul>
&lt;p>&lt;strong>So, I hope I’ve provided a short and precise summary of the three key aspects to focus on when securing your containers. From the first step of building a secure container with a well-written Dockerfile, scanning it during the build process, securing the orchestration platform, to implementing runtime security to monitor and prevent threats in real time. I hope you find this helpful and that it guides you in securing your containerized environments.&lt;/strong>&lt;/p></description></item><item><title>Uncovering the Risks of Third-Party Software Dependencies</title><link>https://benbenhemo.com/post/securing-dependencies/</link><pubDate>Sun, 23 Jun 2024 00:00:00 +0000</pubDate><guid>https://benbenhemo.com/post/securing-dependencies/</guid><description>&lt;h2 id="introduction">&lt;strong>Introduction&lt;/strong>&lt;/h2>
&lt;p>&lt;strong>What comes to mind when you hear the term &amp;ldquo;Software&amp;rdquo;? Do you imagine massive lines of code and complicated algorithms?&lt;/strong>&lt;/p>
&lt;p>Actually, at its core, software is built from various components. These components can include things like snippets of code, or external software modules that add specific functionalities. These external modules, known as dependencies, play a crucial role in modern software development. They allow developers to use existing solutions and speed up their projects.&lt;/p>
&lt;p>In this blog, we&amp;rsquo;ll explore the world of third-party dependencies and the risks associated with them. Understanding and effectively managing these dependencies is crucial for maintaining the security and reliability of applications.&lt;/p>
&lt;p>&lt;strong>Four Key Terms&lt;/strong>&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Third-Party Library&lt;/strong>: External software components used to add functionality without building from scratch.&lt;/li>
&lt;li>&lt;strong>Software Package (or simply package)&lt;/strong>: An library metadata containing the release version of a library, which is a piece of software.&lt;/li>
&lt;li>&lt;strong>Open Source Packages&lt;/strong>: Code made publicly available under an open-source license, allowing for code reviews, community collaboration, and easy reuse in projects.&lt;/li>
&lt;li>&lt;strong>Dependency&lt;/strong>: You probably heard the term before. When you use a specific package in a project, that project “depends” on this package, hence it is called a dependency.&lt;/li>
&lt;/ul>
&lt;h3 id="real-world-example-using-pandas-for-data-analysis">&lt;strong>Real-World Example: Using Pandas for Data Analysis&lt;/strong>&lt;/h3>
&lt;p>Let’s say you’re working on data analysis and you need to process CSV files. To achieve this, you could write numerous functions from scratch, which might involve complex tasks like reading the CSV file, handling missing values, filtering data, and performing calculations. Or you could use an external &lt;strong>third-party library&lt;/strong> called &lt;strong>Pandas&lt;/strong>.&lt;/p>
&lt;ul>
&lt;li>
&lt;p>When you decide to use Pandas, you pull a specific &lt;strong>software package&lt;/strong> of Pandas based on the version that suits your needs. Each version offers different capabilities and improvements.&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" class="chroma">&lt;code class="language-bash" data-lang="bash">&lt;span class="line">&lt;span class="cl">&lt;span class="c1"># Install a specific version of Pandas&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">pip install &lt;span class="nv">pandas&lt;/span>&lt;span class="o">==&lt;/span>1.3.3
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;/li>
&lt;li>
&lt;p>Pandas itself is an &lt;strong>open source package&lt;/strong>, meaning its code is publicly available for anyone to use and contribute to. See: &lt;a href="https://github.com/pandas-dev/pandas" target="_blank" rel="noopener">https://github.com/pandas-dev/pandas&lt;/a>&lt;/p>
&lt;/li>
&lt;li>
&lt;p>When you incorporate Pandas into your project, your project becomes &lt;strong>dependent&lt;/strong> on Pandas, relying on this external library to function correctly. That&amp;rsquo;s why Pandas is part of your software dependencies.&lt;/p>
&lt;/li>
&lt;/ul>
&lt;h2 id="direct-vs-indirect-dependency">Direct vs Indirect Dependency&lt;/h2>
&lt;p>It’s important to understand the difference between direct and indirect dependencies to better assess the risks associated with third-party libraries.&lt;/p>
&lt;h3 id="direct-dependencies">&lt;strong>Direct Dependencies&lt;/strong>&lt;/h3>
&lt;p>A direct dependency is a package that is directly included in the project. Developers intentionally add these dependencies and reference them directly in the code.&lt;/p>
&lt;div class="flex px-4 py-3 rounded-md bg-primary-100 dark:bg-primary-900">
&lt;span class="pr-3 pt-1 text-primary-400">
&lt;svg height="24" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24">&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="m11.25 11.25l.041-.02a.75.75 0 0 1 1.063.852l-.708 2.836a.75.75 0 0 0 1.063.853l.041-.021M21 12a9 9 0 1 1-18 0a9 9 0 0 1 18 0m-9-3.75h.008v.008H12z"/>&lt;/svg>
&lt;/span>
&lt;span class="dark:text-neutral-300">For example, in our project, Pandas is a direct dependency.&lt;/span>
&lt;/div>
&lt;h3 id="indirect-transitive-dependencies">&lt;strong>Indirect (Transitive) Dependencies&lt;/strong>&lt;/h3>
&lt;p>A dependency can use other dependencies for its functionality; these are called indirect dependencies. &lt;strong>A recent study found that an average NPM package can have up to 79 indirect dependencies. Do you understand the depth of this?&lt;/strong>&lt;/p>
&lt;p>These types of dependencies are installed along with the direct dependencies, and the developer usually does not have direct control over which transitive packages are installed.&lt;/p>
&lt;p>Here’s an example of what the dependency tree might look like when you use Pandas:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" class="chroma">&lt;code class="language-bash" data-lang="bash">&lt;span class="line">&lt;span class="cl">&lt;span class="nv">pandas&lt;/span>&lt;span class="o">==&lt;/span>1.3.3
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> ├── numpy&amp;gt;&lt;span class="o">=&lt;/span>1.17.3
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> ├── python-dateutil&amp;gt;&lt;span class="o">=&lt;/span>2.7.3
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> └── pytz&amp;gt;&lt;span class="o">=&lt;/span>2017.3
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;div class="flex px-4 py-3 rounded-md bg-primary-100 dark:bg-primary-900">
&lt;span class="pr-3 pt-1 text-primary-400">
&lt;svg height="24" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24">&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="m11.25 11.25l.041-.02a.75.75 0 0 1 1.063.852l-.708 2.836a.75.75 0 0 0 1.063.853l.041-.021M21 12a9 9 0 1 1-18 0a9 9 0 0 1 18 0m-9-3.75h.008v.008H12z"/>&lt;/svg>
&lt;/span>
&lt;span class="dark:text-neutral-300">In this example, Pandas is a direct dependency, while numpy, python-dateutil, and pytz are indirect dependencies because Pandas relies on them to function properly.&lt;/span>
&lt;/div>
&lt;h2 id="package-management">Package Management&lt;/h2>
&lt;p>You might have heard about the recent news of &lt;a href="https://checkmarx.com/blog/pypi-is-under-attack-project-creation-and-user-registration-suspended/" target="_blank" rel="noopener">PyPi being under attack&lt;/a>, but what exactly is PyPi? PyPi is a trusted Python &amp;ldquo;Package Repository&amp;rdquo; that is widely used by the community.&lt;/p>
&lt;p>A &lt;strong>package repository&lt;/strong> is a centralized location that stores packages, primarily for a specific programming language. The aim of the package repository is to distribute packages more efficiently, providing important information such as metadata, versions, licenses, and indirect dependencies.&lt;/p>
&lt;p>It helps us better understand the packages we plan to use and enhances security and usability by scanning for known vulnerabilities and malware.&lt;/p>
&lt;p>In our case, you can install pandas package from the &lt;a href="https://pypi.org/project/pandas/" target="_blank" rel="noopener">PyPI Repo&lt;/a>. You can use “pip” to install the package from PyPI:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" class="chroma">&lt;code class="language-bash" data-lang="bash">&lt;span class="line">&lt;span class="cl">pip install pandas
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>Package managers are tools that automate the process of installing, upgrading, configuring, and removing software packages in a consistent manner. They are responsible for resolving dependencies and retrieving packages from their respective repositories. Well-known package managers include &lt;strong>pip&lt;/strong> for Python, &lt;strong>npm&lt;/strong> for Node.js, and &lt;strong>Maven&lt;/strong> for Java.&lt;/p>
&lt;h2 id="a-grain-of-pessimism">A Grain Of Pessimism&lt;/h2>
&lt;h3 id="dependency-hell-">&lt;strong>Dependency Hell&lt;/strong> 👹&lt;/h3>
&lt;p>As a developer, you often bear the responsibility for a particular service, which includes developing and maintaining an existing software service. As you may already know, this service can depend on numerous packages and third-party libraries.&lt;/p>
&lt;p>Managing a multitude of dependencies that contain numerous indirect dependencies can be overwhelming. This unmanageable scenario can lead to what is commonly known as &lt;strong>Dependency Hell&lt;/strong>.&lt;/p>
&lt;p>As stated by Wikipedia:
&lt;figure id="figure-wikipedia">
&lt;div class="flex justify-center ">
&lt;div class="w-100" >&lt;img alt="Wikipedia" srcset="
/post/securing-dependencies/wiki_hu3a03657bea4b4d6e0f8876ad1b17b2c5_216470_6bd7a73cf2a520807a215226aff8e59e.webp 400w,
/post/securing-dependencies/wiki_hu3a03657bea4b4d6e0f8876ad1b17b2c5_216470_0b029365bff5d47a3c1a5b75934d8c13.webp 760w,
/post/securing-dependencies/wiki_hu3a03657bea4b4d6e0f8876ad1b17b2c5_216470_1200x1200_fit_q95_h2_lanczos.webp 1200w"
src="https://benbenhemo.com/post/securing-dependencies/wiki_hu3a03657bea4b4d6e0f8876ad1b17b2c5_216470_6bd7a73cf2a520807a215226aff8e59e.webp"
width="760"
height="129"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;figcaption>
Wikipedia
&lt;/figcaption>&lt;/figure>
&lt;/p>
&lt;h3 id="challenges-of-managing-open-source-dependencies">Challenges of Managing Open Source Dependencies&lt;/h3>
&lt;ol>
&lt;li>
&lt;p>&lt;strong>Developers often do not examine the actual code of open-source dependencies&lt;/strong>. If you utilize code reviews in your SDLC process, why not review code from external authors as well? This implicit trust approach can significantly risk your application’s security.&lt;/p>
&lt;/li>
&lt;li>
&lt;p>&lt;strong>Unused Dependencies&lt;/strong>: Changing the package manager file is an easy task that typically requires a merge request process or even a direct commit sometimes. This ease of modification leads developers to add dependencies they think they need for specific functions, but they often end up not using the functionality of these dependencies while still installing them in their application.&lt;/p>
&lt;/li>
&lt;li>
&lt;p>As you have seen, dependencies function like actual software, meaning they are constantly updated and new versions are published to add or modify functionality and apply security patches. &lt;strong>This process puts developers in a problematic situation: they may avoid updating dependency versions due to the possibility of breaking changes, but by not updating, they risk using vulnerable and unpatched dependencies.&lt;/strong>&lt;/p>
&lt;/li>
&lt;/ol>
&lt;h3 id="managing-dependencies-current-security-challenges">Managing Dependencies: Current Security Challenges&lt;/h3>
&lt;h4 id="supply-chain-risks">Supply Chain Risks&lt;/h4>
&lt;p>As you already understand, a single vulnerable dependency can have a massive impact. &lt;strong>All it takes is a single vulnerable piece of code in a &amp;lsquo;hidden&amp;rsquo; dependency used by other popular dependencies to put a wide range of companies at risk.&lt;/strong>&lt;/p>
&lt;h4 id="typosquatting">Typosquatting&lt;/h4>
&lt;p>&lt;strong>Anyone can create an open-source dependency, so why wouldn&amp;rsquo;t attackers?&lt;/strong> Creating a malicious package is easy, and naming it similarly to a known existing package, hoping a victim will misspell the name, is a well-known attack called typosquatting.&lt;/p>
&lt;div class="flex px-4 py-3 rounded-md bg-primary-100 dark:bg-primary-900">
&lt;span class="pr-3 pt-1 text-primary-400">
&lt;svg height="24" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24">&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="m11.25 11.25l.041-.02a.75.75 0 0 1 1.063.852l-.708 2.836a.75.75 0 0 0 1.063.853l.041-.021M21 12a9 9 0 1 1-18 0a9 9 0 0 1 18 0m-9-3.75h.008v.008H12z"/>&lt;/svg>
&lt;/span>
&lt;span class="dark:text-neutral-300">A &lt;a href="https://blog.checkpoint.com/securing-the-cloud/pypi-inundated-by-malicious-typosquatting-campaign/" target="_blank" rel="noopener">recent attack on the PyPi registry&lt;/a> was a classic example of typosquatting.&lt;/span>
&lt;/div>
&lt;h4 id="it-always-about-priorities">It Always About Priorities&lt;/h4>
&lt;p>&lt;strong>Developers are not primarily focused on security. Their main priority is to develop new features that bring value to the business.&lt;/strong> Sometimes, the hard and continuous task of updating dependencies is not something developers are happy to do. Updating to a new version requires them to learn about the vulnerable package and the new updates, and understand how these changes will affect the current software without causing production crashes.&lt;/p>
&lt;h2 id="a-grain-of-optimism">A Grain Of Optimism&lt;/h2>
&lt;p>As you can understand, managing and securing software can be very hard and overwhelming due to the complexity and depth of triaging. We will never be 100% secure, but we can take a few steps to decrease our risk exposure.&lt;/p>
&lt;h3 id="software-bill-of-materials-sbom">Software Bill of Materials (SBOM)&lt;/h3>
&lt;p>With the increase in supply chain attacks, it is crucial for developers to be aware of all dependencies in their projects. Manually cataloging these dependencies is both inefficient and prone to errors.
A Software Bill of Materials (SBOM) provides an automated, machine-readable inventory of all packages and their versions within a project. This helps in better managing and securing dependencies. For more information, visit &lt;a href="https://docs.github.com/en/code-security/supply-chain-security/understanding-your-software-supply-chain/exporting-a-software-bill-of-materials-for-your-repository" target="_blank" rel="noopener">GitHub Docs&lt;/a>.&lt;/p>
&lt;h3 id="detecting-vulnerabilities">Detecting Vulnerabilities&lt;/h3>
&lt;p>Before updating and fixing vulnerabilities in dependencies, you need to detect them right? 🙂 There are two primary approaches to doing this:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Scanner Job in CI/CD Pipelines:&lt;/strong> You can run a dependency scanner as part of your CI/CD pipelines. The scanner compares the versions of your dependencies with a vulnerabilities database and returns a list of detected vulnerabilities that you can start working on. An example of a scanner is &lt;a href="https://github.com/dependabot" target="_blank" rel="noopener">Dependabot&lt;/a>.&lt;/li>
&lt;li>&lt;strong>Manual Scan:&lt;/strong> You can scan your repository manually at a frequency you deem appropriate. The scanner will return the relevant vulnerabilities as well. An example of this is &lt;a href="https://github.com/jeremylong/DependencyCheck" target="_blank" rel="noopener">OWASP Dependency-Check&lt;/a>.&lt;/li>
&lt;/ul>
&lt;h3 id="risk-prioritizations">Risk Prioritizations&lt;/h3>
&lt;p>Each vulnerability has a different risk score. Most of the time, you’ll get the CVSS risk score from the dependency scanners. However, you should prioritize vulnerabilities based on several other factors, such as the specific service’s goal, the presence of sensitive data, public exposure, and more.&lt;/p>
&lt;h3 id="remediation">Remediation&lt;/h3>
&lt;p>Remediation and fixing vulnerabilities in dependencies typically involve updating the dependency to the latest secure version. This should be done in close collaboration with the development departments to ensure nothing breaks during the update.&lt;/p>
&lt;p>You can also check: &lt;a href="https://www.mend.io/blog/the-risks-and-benefits-of-updating-dependencies/" target="_blank" rel="noopener">The Risks and Benefits of Updating Dependencies&lt;/a>&lt;/p>
&lt;p>&lt;strong>Hope this post provided valuable insights into securing third-party software dependencies. Thank you for taking the time to read!&lt;/strong>&lt;/p></description></item><item><title>BigQuery Security 101 - Core Concepts</title><link>https://benbenhemo.com/post/bigquery-security-101/</link><pubDate>Wed, 15 May 2024 00:00:00 +0000</pubDate><guid>https://benbenhemo.com/post/bigquery-security-101/</guid><description>&lt;h1 id="introduction">Introduction&lt;/h1>
&lt;p>Today, in the cloud computing age, the Cloud Service Providers provide services that essentially cut across every aspect of computing and data management needs. From scalable computing power and robust data storage solutions to advanced data analytics and machine learning platforms, these services are designed to empower businesses and developers. They provide the flexibility, scalability, and efficiency required to drive innovation and optimize operations in a digitally transformed world.&lt;/p>
&lt;p>In a brief, these services can broadly be categorized into several types:&lt;/p>
&lt;table>
&lt;thead>
&lt;tr>
&lt;th>Service Category&lt;/th>
&lt;th>Description&lt;/th>
&lt;th>Examples&lt;/th>
&lt;/tr>
&lt;/thead>
&lt;tbody>
&lt;tr>
&lt;td>Compute Services 💻&lt;/td>
&lt;td>Provide virtualized computing resources over the Internet.&lt;/td>
&lt;td>Amazon EC2, Microsoft Azure Virtual Machines, Google Compute Engine&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Storage Services 🗄️&lt;/td>
&lt;td>Dedicated to storing data in the cloud, ensuring security and accessibility.&lt;/td>
&lt;td>Amazon S3, Azure Blob Storage, Google Cloud Storage&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Databases 📚&lt;/td>
&lt;td>Offer scalable, distributed systems for data storage and management.&lt;/td>
&lt;td>Amazon RDS, Azure SQL Database, Google Cloud SQL&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Data Analytics 🔍&lt;/td>
&lt;td>Designed to process and analyze large datasets efficiently.&lt;/td>
&lt;td>AWS Redshift, Azure Synapse Analytics, Google BigQuery&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Machine Learning and AI 🤖&lt;/td>
&lt;td>Enable building, training, and deploying machine learning models.&lt;/td>
&lt;td>AWS SageMaker, Azure Machine Learning, Google AI Platform&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Networking 🌐&lt;/td>
&lt;td>Provide interconnectivity between cloud services, on-premises data centers, and end-users.&lt;/td>
&lt;td>AWS VPC, Azure Virtual Network, Google Cloud VPC&lt;/td>
&lt;/tr>
&lt;/tbody>
&lt;/table>
&lt;p>Today, we’ll zero in on BigQuery, spotlighting the risks that tag along.&lt;/p>
&lt;h1 id="what-is-bigquery">What is BigQuery?&lt;/h1>
&lt;p>💡 &lt;strong>BigQuery, offered by Google Cloud Platform (GCP), stands out as a premier, fully managed, and serverless data warehouse within the Data Analytics spectrum. It enables scalable analysis across petabytes of data, boasting a robust, fast, and interactive SQL-on-terabyte-class database infrastructure. This technology paves the way for the modernization of data-driven applications, ensuring efficient data management and analysis at scale.&lt;/strong>&lt;/p>
&lt;p>For a deeper dive and more context, check out Google&amp;rsquo;s own explanation in the video below:&lt;/p>
&lt;div style="position: relative; padding-bottom: 56.25%; height: 0; overflow: hidden;">
&lt;iframe src="https://www.youtube.com/embed/d3MDxC_iuaw" style="position: absolute; top: 0; left: 0; width: 100%; height: 100%; border:0;" allowfullscreen title="YouTube Video">&lt;/iframe>
&lt;/div>
&lt;h2 id="essential-bigquery-concepts-explained">Essential BigQuery Concepts Explained&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://cloud.google.com/bigquery/docs/jobs-overview" target="_blank" rel="noopener">BigQuery Jobs:&lt;/a> Jobs in BigQuery are specific operations that you can perform on the data stored within. These include executing queries to analyze data, loading new data into tables, or exporting data to different formats or external locations.&lt;/li>
&lt;li>&lt;a href="https://cloud.google.com/bigquery/docs/tables-intro" target="_blank" rel="noopener">BigQuery Tables:&lt;/a> Tables are the fundamental building blocks of BigQuery where your actual data resides. Organized in rows and columns, tables support the storage and retrieval of large quantities of structured data.&lt;/li>
&lt;li>&lt;a href="https://cloud.google.com/bigquery/docs/datasets-intro" target="_blank" rel="noopener">BigQuery Datasets&lt;/a>: A dataset in BigQuery is a container that holds tables and views. Datasets are used to organize and control access to your tables based on needs and are defined at the project level.&lt;/li>
&lt;li>&lt;a href="https://cloud.google.com/bigquery/docs/reference/standard-sql/query-syntax" target="_blank" rel="noopener">BigQuery Queries:&lt;/a> Queries in BigQuery are used to interact with the data stored in tables. By writing SQL-like commands, you can perform complex data analysis, manipulate data, and generate insights.&lt;/li>
&lt;/ul>
&lt;h1 id="now-lets-talk-security">Now Let’s Talk Security&lt;/h1>
&lt;p>Supporting companies with huge data stores—and some may be containing sensitive information, such as PII, financial records, or intellectual property—makes BigQuery capabilities for advanced data analytics more than mandatory. Such an approach equalizes to an increased level of sensitivity and, accordingly, turns BigQuery into a priority target of cyber threats. Attackers will exploit this service for carrying out such actions as access and exfiltration of sensitive data but not limited to DoS or misuse of computational resources.&lt;/p>
&lt;h2 id="unauthorized-access">&lt;strong>Unauthorized Access&lt;/strong>&lt;/h2>
&lt;blockquote>
&lt;p>&lt;strong>The Problem →&lt;/strong> Gaining unauthorized access to BigQuery can lead to sensitive data exposure.&lt;/p>
&lt;/blockquote>
&lt;blockquote>
&lt;p>&lt;strong>The Solution →&lt;/strong> Adopt the principle of Least Privilege Access, granting users only the access they need.&lt;/p>
&lt;/blockquote>
&lt;h3 id="1the-iam-solution">1️⃣ &lt;strong>The IAM Solution&lt;/strong>&lt;/h3>
&lt;p>You might think everyone knows about IAM (Identity and Access Management) by now, but getting it right is crucial for keeping your BigQuery data safe.
&lt;a href="https://cloud.google.com/bigquery/docs/access-control" target="_blank" rel="noopener">BigQuery Access control with IAM&lt;/a> provides detailed guidance on establishing strict IAM policies.&lt;/p>
&lt;p>I&amp;rsquo;ve included an overview of the roles, just in case:&lt;/p>
&lt;p>&lt;strong>High-Risk BigQuery Roles&lt;/strong>&lt;/p>
&lt;table>
&lt;thead>
&lt;tr>
&lt;th>Role&lt;/th>
&lt;th>Permissions&lt;/th>
&lt;th>Security Risk&lt;/th>
&lt;/tr>
&lt;/thead>
&lt;tbody>
&lt;tr>
&lt;td>BigQuery Admin (&lt;code>roles/bigquery.admin&lt;/code>)&lt;/td>
&lt;td>Full control over all BigQuery resources.&lt;/td>
&lt;td>Can lead to data exposure, loss, or unauthorized modification.&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>BigQuery Data Owner (&lt;code>roles/bigquery.dataOwner&lt;/code>)&lt;/td>
&lt;td>Full control over datasets and their contents.&lt;/td>
&lt;td>Risk of unauthorized access or modifications within datasets.&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>BigQuery User (&lt;code>roles/bigquery.user&lt;/code>)&lt;/td>
&lt;td>Create new datasets and run jobs.&lt;/td>
&lt;td>Potential for unauthorized dataset creation or costly/malicious job execution.&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>BigQuery Job User (&lt;code>roles/bigquery.jobUser&lt;/code>)&lt;/td>
&lt;td>Run jobs, including queries, within the project.&lt;/td>
&lt;td>May be exploited for data exfiltration or incurring high costs.&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>BigQuery Data Editor (&lt;code>roles/bigquery.dataEditor&lt;/code>)&lt;/td>
&lt;td>Read/write access to datasets, no sharing.&lt;/td>
&lt;td>Can lead to data loss or unauthorized alterations.&lt;/td>
&lt;/tr>
&lt;/tbody>
&lt;/table>
&lt;ul>
&lt;li>For more details on BigQuery permissions and API calls, visit this page: &lt;a href="https://gcp.permissions.cloud/iam/bigquery" target="_blank" rel="noopener">https://gcp.permissions.cloud/iam/bigquery&lt;/a>&lt;/li>
&lt;/ul>
&lt;p>&lt;strong>allUsers/allAuthenticatedUsers&lt;/strong>&lt;/p>
&lt;p>Another important aspect is to ensure that no publicly accessible BigQuery datasets are available within your GCP environment.
Make sure that role bindings such as allUsers and allAuthenticatedUsers are not configured:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>&amp;ldquo;allUsers:&amp;rdquo;&lt;/strong> This allows any user on the internet, whether authenticated or unauthenticated, to access your dataset.&lt;/li>
&lt;li>&lt;strong>&amp;ldquo;allAuthenticatedUsers:&amp;rdquo;&lt;/strong> This allows any user who can sign in to GCP to access your dataset.&lt;/li>
&lt;/ul>
&lt;h3 id="2row-level-security-rls">2️⃣ &lt;strong>Row-Level Security (RLS)&lt;/strong>&lt;/h3>
&lt;p>&lt;a href="https://cloud.google.com/bigquery/docs/row-level-security-intro" target="_blank" rel="noopener">Row-Level Security (RLS)&lt;/a> in BigQuery enhances access control, extending it down to the granularity of table rows. You can implement access control at the project, dataset, and table levels, as well as &lt;a href="https://cloud.google.com/bigquery/docs/column-level-security-intro" target="_blank" rel="noopener">column-level security&lt;/a> through policy tags. This method allows for more delicated access control, by deciding which entity can access specific columns.&lt;/p>
&lt;p>I think Google&amp;rsquo;s explanation regarding &lt;a href="https://cloud.google.com/bigquery/docs/row-level-security-intro" target="_blank" rel="noopener">Row Level Security&lt;/a> is excellent, so I will use their use case to demonstrate how RLS is implemented:&lt;/p>
&lt;ul>
&lt;li>Think of a table, &lt;strong>&lt;code>dataset1.table1&lt;/code>&lt;/strong>, where rows are labeled by different regions in the &lt;strong>&lt;code>region&lt;/code>&lt;/strong> column.&lt;/li>
&lt;li>Row-level security allows a Data Owner or Admin to set specific policies, such as allowing only members of the &lt;strong>&lt;code>group:apac&lt;/code>&lt;/strong> to see data from the APAC region.&lt;/li>
&lt;li>As a result, only users in the &lt;strong>&lt;code>sales-apac@example.com&lt;/code>&lt;/strong> group can view rows where the &lt;strong>&lt;code>Region&lt;/code>&lt;/strong> is &amp;ldquo;APAC&amp;rdquo;. Similarly, those in the &lt;strong>&lt;code>sales-us@example.com&lt;/code>&lt;/strong> group can access rows marked as &amp;ldquo;US&amp;rdquo;. Users who aren&amp;rsquo;t in either the APAC or US groups will not be able to see any rows.&lt;/li>
&lt;li>Under the row-level access policy called &lt;strong>&lt;code>us_filter&lt;/code>&lt;/strong>, various entities, including the chief US salesperson &lt;strong>&lt;code>jon@example.com&lt;/code>&lt;/strong>, are granted access to rows associated with the US region.&lt;/li>
&lt;/ul>
&lt;p>
&lt;figure >
&lt;div class="flex justify-center ">
&lt;div class="w-100" >&lt;img src="https://cloud.google.com/static/bigquery/images/row_level_security_use_case_regions.png" alt="From Google Documentation: The resulting behavior is that users in the group &amp;lt;code&amp;gt;sales-apac@example.com&amp;lt;/code&amp;gt; can view only rows where &amp;lt;code&amp;gt;Region = &amp;amp;quot;APAC&amp;amp;quot;&amp;lt;/code&amp;gt;. Similarly, users in the group &amp;lt;code&amp;gt;sales-us@example.com&amp;lt;/code&amp;gt; can view only rows in the &amp;lt;code&amp;gt;US&amp;lt;/code&amp;gt; region. Users not in &amp;lt;code&amp;gt;APAC&amp;lt;/code&amp;gt; or &amp;lt;code&amp;gt;US&amp;lt;/code&amp;gt; groups don&amp;amp;rsquo;t see any rows." loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>From Google Documentation: The resulting behavior is that users in the group &lt;code>sales-apac@example.com&lt;/code> can view only rows where &lt;code>Region = &amp;quot;APAC&amp;quot;&lt;/code>. Similarly, users in the group &lt;code>sales-us@example.com&lt;/code> can view only rows in the &lt;code>US&lt;/code> region. Users not in &lt;code>APAC&lt;/code> or &lt;code>US&lt;/code> groups don&amp;rsquo;t see any rows.&lt;/p>
&lt;h2 id="data-exfiltration">&lt;strong>Data Exfiltration&lt;/strong>&lt;/h2>
&lt;blockquote>
&lt;p>&lt;strong>The Problem →&lt;/strong> Through unauthorized queries and jobs, attackers can export data to external storage for nefarious purposes&lt;/p>
&lt;/blockquote>
&lt;blockquote>
&lt;p>&lt;strong>The Solution →&lt;/strong> Implementing Robust Monitoring and Threat Detection Strategies&lt;/p>
&lt;/blockquote>
&lt;h3 id="a-word-on-bigquery-log-versions">&lt;strong>A Word On BigQuery Log Versions&lt;/strong>&lt;/h3>
&lt;p>If you are using BigQuery in your company, you may have noticed that BigQuery has two log versions. In short, here is the difference:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>&lt;code>AuditData&lt;/code>&lt;/strong> is the legacy version of BigQuery audit logs, primarily monitoring API calls.&lt;/li>
&lt;li>&lt;strong>&lt;code>BigQueryAuditMetadata&lt;/code>&lt;/strong> is the “v2” version of the logs, similar to AuditData. It monitors the activities in BigQuery, such as executeing jobs and queries, reading and updating tables and datasets.&lt;/li>
&lt;/ul>
&lt;p>We&amp;rsquo;re not going to delve deeply into which detection rules you can create, but just by examining the &lt;strong>&lt;code>protoPayload.methodName&lt;/code>&lt;/strong>, I believe you can come up with a few ideas 🙂 :&lt;/p>
&lt;table>
&lt;thead>
&lt;tr>
&lt;th>Method&lt;/th>
&lt;th>Description&lt;/th>
&lt;/tr>
&lt;/thead>
&lt;tbody>
&lt;tr>
&lt;td>&lt;code>google.cloud.bigquery.v2.TableService.InsertTable&lt;/code>&lt;/td>
&lt;td>Creates a new table.&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>google.cloud.bigquery.v2.TableService.UpdateTable&lt;/code>&lt;/td>
&lt;td>Replaces a table’s metadata.&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>google.cloud.bigquery.v2.TableService.PatchTable&lt;/code>&lt;/td>
&lt;td>Updates parts of a table’s metadata.&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>google.cloud.bigquery.v2.TableService.DeleteTable&lt;/code>&lt;/td>
&lt;td>Removes a table.&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>google.cloud.bigquery.v2.DatasetService.InsertDataset&lt;/code>&lt;/td>
&lt;td>Creates a new dataset.&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>google.cloud.bigquery.v2.DatasetService.UpdateDataset&lt;/code>&lt;/td>
&lt;td>Replaces a dataset’s metadata.&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>google.cloud.bigquery.v2.DatasetService.PatchDataset&lt;/code>&lt;/td>
&lt;td>Updates parts of a dataset’s metadata.&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>google.cloud.bigquery.v2.DatasetService.DeleteDataset&lt;/code>&lt;/td>
&lt;td>Deletes a dataset.&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>google.cloud.bigquery.v2.TableDataService.List&lt;/code>&lt;/td>
&lt;td>Lists table data.&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>google.cloud.bigquery.v2.JobService.InsertJob&lt;/code>&lt;/td>
&lt;td>Submits a processing job.&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>google.cloud.bigquery.v2.JobService.Query&lt;/code>&lt;/td>
&lt;td>Executes a query and returns results.&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>google.cloud.bigquery.v2.JobService.GetQueryResults&lt;/code>&lt;/td>
&lt;td>Retrieves results of a completed query.&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>InternalTableExpired&lt;/code>&lt;/td>
&lt;td>Indicates a table was auto-deleted after expiring.&lt;/td>
&lt;/tr>
&lt;/tbody>
&lt;/table>
&lt;blockquote>
&lt;p>💡 You can also check out this great &lt;a href="https://cloudsecurityalliance.org/blog/2022/11/01/planning-for-attacks-how-to-hunt-for-threats-in-bigquery" target="_blank" rel="noopener">blog post&lt;/a> by Lionel Saposnik and Dan Abramov, which provides actual examples of threat hunting in BigQuery. One of the use cases shows how to hunt for data being exported to an external dataset.&lt;/p>
&lt;/blockquote>
&lt;h2 id="data-masking">Data Masking&lt;/h2>
&lt;blockquote>
&lt;p>&lt;strong>The Problem →&lt;/strong> Sometimes, sensitive data within BigQuery should be accessible to users who have legitimate system access but don&amp;rsquo;t need to see all the details, such as for analysis purposes.&lt;/p>
&lt;/blockquote>
&lt;blockquote>
&lt;p>&lt;strong>The Solution →&lt;/strong> Data masking in BigQuery provides a solution by concealing specific data elements, allowing users view only to the information essential for their roles.&lt;/p>
&lt;/blockquote>
&lt;h3 id="benefits-of-data-masking">&lt;strong>Benefits of Data Masking&lt;/strong>&lt;/h3>
&lt;p>Data masking offers several important benefits that enhance both security and operational efficiency:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Streamlines Data Sharing&lt;/strong>: By masking sensitive columns, you can safely share tables with larger groups without compromising sensitive information.&lt;/li>
&lt;li>&lt;strong>Maintains Query Integrity&lt;/strong>: Data masking works seamlessly with existing queries. Configuring data masking ensures that sensitive data is automatically obscured, based on the roles assigned to users, without the need to modify each query.&lt;/li>
&lt;li>&lt;strong>Scalable Data Policies&lt;/strong>: You can establish a data policy, associate it with a policy tag, and apply this tag across numerous columns. This approach allows for the consistent application of access rules on a wide scale.&lt;/li>
&lt;li>&lt;strong>Enables Attribute-Based Access Control&lt;/strong>: By attaching a policy tag to a column, data access becomes contextual, governed by the specifics of the data policy and the roles associated with that policy tag. This method ensures data is accessible only under appropriate circumstances.&lt;/li>
&lt;/ul>
&lt;h3 id="configuring-data-masking">Configuring Data Masking&lt;/h3>
&lt;p>To set up data masking in BigQuery, I suggest following the straightforward guide provided by Google Cloud. It covers everything you need to know to get started: &lt;a href="https://cloud.google.com/bigquery/docs/column-data-masking-intro" target="_blank" rel="noopener">Google Cloud Guide to Data Masking in BigQuery&lt;/a>.&lt;/p>
&lt;p>In summary, here&amp;rsquo;s what you need to do:&lt;/p>
&lt;ol>
&lt;li>&lt;strong>Set up a taxonomy with policy tags&lt;/strong>: A taxonomy is a framework you design to categorize your data. Within this framework, policy tags are the markers you&amp;rsquo;ll use to indicate which data is sensitive.&lt;/li>
&lt;li>&lt;strong>Create data policies for policy tags&lt;/strong>: These are the rules attached to your policy tags. Data policies determine who can see data unmasked and who sees it masked.&lt;/li>
&lt;li>&lt;strong>Set policy tags on columns&lt;/strong>: Once you&amp;rsquo;ve got your taxonomy and policies in place, you&amp;rsquo;ll assign the policy tags to the specific columns in your BigQuery tables that contain sensitive data.&lt;/li>
&lt;li>&lt;strong>Grant access through the Masked Reader role&lt;/strong>: This role is for users who should see the masked data. By assigning users to the Masked Reader role at the data policy level, they&amp;rsquo;ll only be able to access data according to the policy you&amp;rsquo;ve set.&lt;/li>
&lt;/ol>
&lt;p>
&lt;figure >
&lt;div class="flex justify-center ">
&lt;div class="w-100" >&lt;img src="https://cloud.google.com/static/bigquery/images/data-masking-workflow.png" alt="https://cloud.google.com/static/bigquery/images/data-masking-workflow.png" loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>&lt;strong>Using BigQuery?&lt;/strong> I hope this post has given you a solid introduction to what BigQuery can do and how to protect it. If you have any insights or stories about BigQuery security, please share them in the comments&lt;/p></description></item></channel></rss>