Practitioner notes from r/dataengineering and r/databricks

all clouds2026-10-037 min read

Top/month posts from the combined subreddit RSS feed. The list keeps all r/databricks posts. It keeps r/dataengineering posts only when the title or summary names Databricks. The feed preserves Reddit's ranking but does not expose scores or comment counts.

This is a dated reading list, not product guidance. Verify technical claims against official documentation.

Databricks and lakehouse#

  • WHAT IS DATABRICKS? — r/databricks · 2026-09-04 Let's pretend someone knew nothing about Databricks. How would you explain it? This is something I have been asked multiple times throughout my career, curious to hear how others would explain Databricks to a complete beginner. In my opinion, Databricks is basically a place where companies put all their data so they
  • Databricks SSH Tunnel for connecting your coding agents and IDEs to your workspace — r/databricks · 2026-09-03 Just made a video about a feature I'm pretty excited about. tldr: You can use an SSH tunnel to connect your coding agents and IDE (VSCode/Cursor) to your Databricks workspace. See the video for a full walkthrough. Some notes on things I forgot to mention in the video: - claude/codex isn't natively installed when you
  • What frustrates you when using Databricks? — r/databricks · 2026-09-03 Any common bugs, features you would like to see, or underrated useful features more people should know about?
  • Personal Opinion about AI Taking Over Data Engineering Roles and Impact on New Devs Entering The Market — r/dataengineering · 2026-09-05 Fresh grad doing a DE role with a small company. We do not use unified cloud data platforms like Databricks/Snowflake, so the entire tech stack for the data pipeline is selectively chosen, though some of the selections are cloud hosted -> think kafka, flink, columnar DB, row-based DB, airflow for orchestration.
  • How to make it more near real time — r/dataengineering · 2026-09-05 We have a Databricks pipeline where we receive source JSON messages through Azure Service Bus. A job polls the Service Bus every 20 mins during business hours, downloads the messages and fans out the data into around 82 Bronze tables. The writes are not always simple appends. For some tables we to MERGE. One batch
  • Is the Bronze → Silver → Gold architecture still the best approach for every Databricks project? — r/databricks · 2026-09-03 I’ve been learning about the Medallion Architecture in Databricks, where data typically moves through Bronze, Silver, and Gold layers. It makes sense for many data platforms, but I’m curious about real-world implementations. Do you think Bronze → Silver → Gold is still the best approach for every Databricks project?
  • What’s one Databricks “best practice” you disagree with? — r/databricks · 2026-09-03 Something that sounds great in Databricks documentation but didn't make sense for your workload in production? Curious what people have learned the hard way.
  • Are we over-optimizing Delta tables? — r/databricks · 2026-09-03 Between OPTIMIZE, Z-ORDER, liquid clustering, partitioning, and automatic optimization, it sometimes feels like we're spending more time optimizing tables than querying them. How do you decide which optimizations are actually worth it in production?
  • WLB Databricks GTM — r/databricks · 2026-09-07 Considering a GTM role with Databricks. Compelling role, comp etc., but cannot get a proper read on the WLB and culture. Have a little one at home, can’t afford a job that requires major travel or 12+ hrs work. Love to hear from people in the company on what the reality on the ground is like.
  • How data engineer do effective testing in Databricks? — r/databricks · 2026-09-05 I have been writing SQL scripts to ensure data sanity.What are the other ways? Is pytest useful? Let's say, i populated my bronze table from the source. I want to check if the correct mapping is done. I wrote SQL scripts. What are better ways
  • Dynamic Select is now in Lakeflow Designer — r/databricks · 2026-09-03 You can now dynamically select columns in Lakeflow Designer. This makes it easy to bulk-keep, or bulk-drop columns from a very wide table. And your data prep will keep working as your underlying schema evolves.
  • Query works in Databricks SQL but fails through JDBC — r/databricks · 2026-09-03 I’ve noticed cases where a query runs successfully in the Databricks SQL editor but fails when the exact same query is executed through a JDBC-based application or data quality tool. Has anyone run into this? Was the issue related to query wrapping, session settings, SQL dialect differences, or JDBC driver behavior?
  • Databricks Production Planning: How to Actually Use the Deployment Guide — r/databricks · 2026-09-07 A practical read of the 10-phase Databricks deployment guide: what to decide first, and how to use it on a platform you already run. Most Databricks platforms get designed one of two ways. On the fly, project by project, as teams onboard and workspaces appear. Or properly, once, right at the start, and then never
  • Lakeflow Genie Code Task — r/databricks · 2026-09-03 We now have Genie Code Task in Lakeflow jobs. Can not yet send output to if/else, but more options for orchestration are planned. more news
  • Community BrickTalk | One Platform, Any Source: Unifying Enterprise Data with Lakeflow Connect — r/databricks · 2026-09-07 Hey r/Databricks! We’re hosting a free, community-sponsored BrickTalk on Thursday, September 17, 2026, focusing on how to simplify and scale data ingestion using Lakeflow Connect! BrickTalks is a community event series where Databricks experts share real-world use cases, live demos, and practical insights, giving you
  • Apache Iceberg Compaction Best Practices — r/databricks · 2026-09-06 submitted by /u/codingdecently to r/databricks [link] [comments]
  • How to do cross-cloud sharing with OpenSharing with added security (demo) — r/databricks · 2026-09-04 Hey folks! In this demo, Akram from Databricks' product team shares how you can leverage SecureConnect to better your security posture when doing cross-cloud sharing on OpenSharing! If you have no idea what OpenSharing is, how it applies to you, or how we got from Delta Sharing to OpenSharing, also encourage you to
  • Automating Apache Iceberg Table Maintenance — r/databricks · 2026-09-06 submitted by /u/codingdecently to r/databricks [link] [comments]
  • Acquiring/Processing from a MQ to a Delta — r/databricks · 2026-09-06 Has anyone tried acquiring data from a MQ at scale using apache spark on databricks cluster? I was trying to solve this problem at work but so far havn't seen an native lib or efficient ways to do this. The legacy system seems to be pulling data using a java based utility and wanted to if there are any imporvements or
  • I merged two databases (Postgres and Elasticsearch) into Lakebase, then threw 200 AI agents at it. — r/databricks · 2026-09-07 submitted by /u/Limp-Park7849 to r/databricks [link] [comments]

Other practitioner signal#

  • What are the books on Matei's shelf? — r/databricks · 2026-09-03 As a techie, I always wonder what the founders are reading. The only title that's visible is Think Lego Bricks. But I can't find the title on Amazon. What about the others? The one on the right of Lego Bricks is a tech book. And there's an O'Reilly book. Maybe the book he wrote? Spark: The Definitive Guide. But it
  • [Private Preview] Concurrent Write Support for Identity Columns! — r/databricks · 2026-09-03 What are concurrent identity columns? A new implementation of identity columns that supports concurrent writes. You can use this query to find the tables with the most amount of concurrent transaction failures due to identity columns. How to enable CREATE TABLE new_identity_table (id BIGINT GENERATED ALWAYS AS
  • Lakebase: Serverless Postgres over Open Lake Storage — r/databricks · 2026-09-04 Interesting paper to read about Lakebase.
  • Omnigent Local Coding Model Rec — r/databricks · 2026-09-06 After watching Matei's webinar and the post on controlling spend, been trying to use the other harnesses and models folks are suggest and trying out Qwen 2.5 coding and 3.6 with Polly in local Omnigent (not connected to a workspace). I have Codex and Claude but ideally thinking best to use paid higher model to plan
  • Serverless Env v6 — r/databricks · 2026-09-05 Version 6 of the serverless environment is available, which corresponds to runtime 19. more news