-

The Real Reason Data Engineers Fail Technical Interviews
You can be good at SQL and still struggle in a Data Engineer interview. You can work with Snowflake every day. You can build production pipelines, write dbt models, debug data issues and understand your company’s entire data platform. Then the interviewer gives you a problem you’ve technically seen before: “Find users who have logged…
-

Snowflake DCM Projects: Infrastructure as Code, Native
Until August 2026, managing Snowflake infrastructure as code meant one of two paths. You either adopted Terraform with the Snowflake provider — an external tool with its own state file, version lag behind new Snowflake features, and a separate CI/CD pipeline to learn and maintain. Or you ran Schemachange or a migration-script pattern — imperative…
-

2extract for Web Scraping: Python, Proxies & MCP Guide
A ten-thousand-row scraping job died two hours in, wallet already charged for the first six thousand rows, when a regional storefront started quietly returning empty product grids instead of a clean 403 I could at least catch. A request from a cloud IP gets fingerprinted as automation by anything that matters — pricing pages, ad…
-

dbt State: Stop Rebuilding What Hasn’t Changed
Every hour, on a schedule, your dbt job rebuilds every model in your project. The staging models run because they always run. The intermediate joins run because they always run. The mart that powers the executive dashboard runs because, well, it always runs — even when not a single row in any upstream table changed…
-

Running LLM Tasks in Apache Airflow with the common.ai Provider
For the past two years, the standard pattern for running LLM calls in Airflow was a PythonOperator that imported the OpenAI client, called the API, and returned the result as XCom. It worked. But when the call failed at 3 AM, Airflow retried the entire task — including the database query that produced the input.…
-

Snowflake Dynamic Data Masking & Row Access Policies: A Production Guide
The governance review lands on a Wednesday. Your company needs to prove that analysts in one region cannot see customer PII from another, that customer emails are masked for anyone below the data steward tier, and that — this one is new — AI agents querying your warehouse cannot extract raw PII even when the…
-

Building Data Pipelines That Feed AI Features Without Breaking the Bill
When the consumer at the end of a pipeline is a language model rather than a dashboard, the expensive step moves from the transform to the last hop, and every row you push through it costs money. This article covers the pipeline shape we keep returning to, a four-question test for choosing batch, event-driven or…
-

Snowflake Cortex AI Token Usage Monitoring: The Complete Guide
Somewhere on your team, an AI_CLASSIFY job is running on a table larger than anyone realised. Or a Cortex Agent is looping through a multi-step workflow that seemed cheap in testing. Or a developer left a search service indexed and running in a dev environment that nobody is querying anymore. None of these will trigger…
-

Using MCP Servers with Snowflake: A Practitioner’s Guide
Your data team ships a Cortex-powered analytics agent. Works beautifully. Then the platform team wants to plug in Cursor. The ML team asks about GPT-4o. A product manager hears about Claude Desktop and sends a Slack message. Suddenly you’re the person maintaining four different Snowflake connectors, each with its own auth token, its own privilege…
-

From ETL Pipelines to Data Products: Designing Reusable Data Infrastructure for Enterprise AI
Enterprise data platforms often begin with a simple objective: move data from operational systems into a place where it can be analyzed. Over time, however, the number of data sources, consumers, and business requirements grows. A pipeline originally created for one dashboard becomes useful to another team. A transformation developed for a reporting workload is…