↓Skip to main content

When to use Databricks

·5 mins · Databricks Tools
Author
Alex

If you like this post, you should subscribe. Or, send me an email directly.

I just learned about Databricks, we should start using it.

My company has signed an enterprise deal with Databricks and now we are all learning it.

Most of us have likely heard about Databricks.

But when should you really use Databricks? You should absolutely use Databricks, but only for the narrow use cases when you absolutely have to.

In this article I go through when and how to use Databricks, so that you can maximize the value of it and minimize its cost. We’ll cover both a tactical data scientist or developer perspective, and a strategic, business perspective.

When to Use Databricks as a Developer #

You should use Databricks when you need distributed processing using Spark, because your data is too large to easily handle on one machine.

When developing data science applications or just doing data science analysis, the highest value compute is the one you are doing yourself as you are iterating on the code, model or analysis. This iteration speed is directly bottlenecked by how quickly you can try an idea and get feedback, i.e. how quickly you can run it.

The most valuable compute is the one you do yourself in your head.

Therefore you should always try to work “in memory”, i.e. in an interactive R or Python session on a single machine, because this gives you almost instantaneous feedback. Just starting a Databricks cluster can take minutes. If you are developing Spark applications, you can do so locally, using sbt ~testQuick to continuously watch for file changes and then run only failing tests. This is strictly speaking not “in memory” like when you develop interactively, but the equivalent for a compiled language where the feedback is given by unit tests.

On a modern machine you can easily work with millions of rows in memory, which is often enough either for an aggregated dataset or a sample. Even when the total available data is large you can often still do lots of work in memory. Just sample, making sure to do it correctly. It’s only really when you need exact counts, like reporting, that you can’t sample. Otherwise application development, model building and analysis can be done on a sample.

The faster your model runs, the more chances to improve it you get.

Just make sure to periodically run your code on the full data. For this you can and should use Databricks.

When you do need to run large scale data processing, it’s again important to think about the cost of your own compute. If you, or the organization you are part of, has the capacity to manage your own lower-level Spark compute, like AWS EMR or Google Cloud Dataproc, then, by all means use that. But, if you find yourself spending too much time just getting jobs to run, it will probably be cheaper to use Databricks.

But even if you do have Databricks available, my recommendation is to be deliberate about when to use it. Otherwise you will spend a lot of time waiting for cluster upsizing etc., which means that your most valuable compute, yourself, will have less time to try new ideas and refinements.

How to Use Databricks as a Developer #

The key to maximizing Databricks value is to treat it as a commoditized compute provider, even if it isn’t. What I mean by this is that you define jobs and clusters in code, then use the CLI to interact with Databricks to create, read, update and delete jobs and clusters. Write scripts on top of the CLI to make this ergonomic. Integrate it into your CI process.

Build a deliberate development loop that iterates quickly locally and in memory, and then validates on full scale data on Databricks. You can even have scheduled, nightly validation jobs that run while you sleep.

When to Use Databricks as a Business #

Databricks’ killer feature is the ability to easily run and schedule Spark jobs for large amounts of data. This is the key feature that you as a business want to buy, unless you happen to have, or want to invest in, the capacity to manage lower-level Spark compute like AWS EMR or Google Cloud Dataproc, and can amortize the cost of this capacity over a large volume of Spark jobs.

Databricks, like any SaaS provider, has a lot of features. You only need, and should only use, a small subset of these. This will help prevent vendor lock-in and make your team more capable and nimble. Use a few building blocks that are the best in class for their use case: for running Spark code, that is Databricks.

Concretely, this means Serverless compute and Classic compute as well as Lakeflow Jobs for scheduling.

Some features don’t incur any additional fees, so sometimes it seems like a good deal to use them, since they are bundled anyway. I still recommend being strategic about the use of those features, because using them can increase the lock-in costs, thus providing less freedom to explore other solutions for the expensive features. For example, if you start relying on Databricks AI/BI dashboards for reporting, that can make it harder to switch vendors for running your expensive Spark jobs.

Conclusion #

Always work locally when you can. That said, if Databricks speeds up your development feedback loop, your most valuable compute, then you should absolutely use it. But be deliberate about how you use it. Pay the premium until you can amortize the cost of a lower-level platform. Once you can, then having stayed disciplined about features will pay off when switching.