Databricks is one of the most capable data and AI platforms available. Morgan Stanley runs its fully managed lakehouse on Databricks. Goldman Sachs uses it for data governance at enterprise scale. It powers AI and ML workloads at some of the most sophisticated financial institutions in the world. That is not a platform to dismiss. The question firms face when evaluating Databricks as the foundation for an investment data platform is a specific one: does your firm have, or can it realistically build and retain, the engineering team capable of doing that work? And is building that team the right use of your capital and talent budget? Databricks is not a shortcut to an investment data platform. It is the most powerful set of tools available to build one from scratch.
Note: Information on Databricks' platform capabilities is based on Databricks' public documentation, product announcements, and BARC research reviewed June 2026.
Databricks is a data intelligence platform built on Apache Spark, Delta Lake, and a growing suite of AI and ML tooling, including Mosaic AI, MLflow for model governance, and Unity Catalog for data governance. It excels at large-scale data engineering, machine learning pipelines, and AI model development and deployment.
Databricks launched Lakeflow Connect in 2025 for investment management use cases, with connectivity to portfolio accounting systems, risk platforms, and fund administration systems. It has a published reference architecture for building a modern investment data platform, covering securities reference data, portfolio analytics, risk calculations, and trade lifecycle management.
To execute on that reference architecture, your firm needs engineers proficient in Apache Spark, Python or Scala, Delta Lake, cluster management, and ML governance. BARC research identifies Databricks as having longer-than-average project delivery timelines compared to peer platforms. That reflects the nature of a platform this powerful and flexible: it requires a serious engineering investment to deliver serious results.
Investment data environments have become more complex. Firms increasingly require consistent data across front, middle, and back-office functions, and the integration scope, counterparties, fund administrators, market data vendors, grows as strategies evolve. At the same time, the talent required to build and maintain a Databricks-based investment data platform is among the most sought-after in financial services technology.
As a result, evaluating a Databricks-based build is less about Databricks' capability as an engineering platform, which is well-established, and more about whether your firm has the engineering depth and time horizon to execute on it, and whether that investment generates the most value relative to alternatives.
Databricks was built by the creators of Apache Spark. Accessing its power requires engineers who understand Spark execution, Delta Lake architecture, cluster optimization, and ML governance, simultaneously with investment domain knowledge. This profile is rare and expensive. Setting up Databricks is routinely described as complex and time-consuming even for technical teams, involving data engineers, machine learning engineers, and other specialists. Scala, the language Spark is written in, offers the best performance but is difficult to hire for. The engineering depth required to build on Databricks effectively is among the highest of any data platform on the market.
Databricks published a reference architecture for investment management in 2025, covering securities reference data, portfolio analytics, risk calculations, trade lifecycle management, and connectivity to Bloomberg, LSEG, Charles River, and Aladdin. It is a well-constructed guide for firms with the engineering capacity to follow it. Working through that architecture in production requires resolving hundreds of investment domain questions it does not answer: how to handle corporate actions edge cases, how to model counterparty reconciliation breaks, how to implement investment-specific data quality rules. Those answers have to be built by your engineering team and maintained as requirements evolve.
BARC research across data platform users identifies Databricks as having longer-than-average project delivery timelines. That reflects a structural characteristic of the platform: it is optimized for power and flexibility, not for rapid deployment. Building a production-grade investment data platform on Databricks is a multi-year undertaking for most investment management firms. Every month spent building rather than operating is a month of operational risk and unrealized value. Purpose-built investment data platforms deliver incremental value from deployment, without requiring a multi-year engineering program first.
Firms are increasingly distinguishing between AI and ML engineering platforms and investment data platforms. Databricks excels at the former. Investment data management, ingestion, normalization, quality, distribution, and analytics across the investment lifecycle, is a different problem, and one that benefits from domain-specific expertise embedded in the product rather than assembled through a custom engineering program.
This distinction matters because the two use cases require different things from the teams using them. Databricks is built for data scientists and engineers building models. Investment data management requires platforms built for the teams running operations, with self-service configuration, no-code tooling, and governance applied consistently across the investment lifecycle.
The most effective architectures often combine both: a governed investment data layer feeding downstream Databricks environments for quantitative and AI workloads where Databricks' power is genuinely differentiated.
Firms evaluating a Databricks-based build are typically weighing engineering cost and capability against speed to operational value.
A Databricks-based investment data build is the right choice for a specific profile: large technology organizations with mature data engineering teams who know Spark, quantitative research operations requiring custom ML model development at scale, and firms with multi-year technology horizons that can absorb the project timelines the platform requires.
For most investment management firms, particularly those focused on operational efficiency, multi-asset data management, or reducing engineering overhead, Aquata provides a more direct path to operational capability without the engineering program Databricks requires.
The most important question is not whether Databricks is a capable engineering platform, it is, and it is genuinely superior for AI and ML model development at scale. The question is whether your firm's primary investment data requirement is building custom ML models or managing investment operations data, and whether the engineering investment required to do both on a single platform is the best use of your resources.
Databricks and Aquata are not mutually exclusive. Several Arcesium clients use Databricks for model development and quantitative research on top of data that Aquata has already normalized and governed. That architecture, Aquata as the governed investment data layer, Databricks downstream for quantitative workloads, combines the operational efficiency of a purpose-built platform with the ML power of Databricks, without requiring your investment operations team to depend on Spark expertise.
Aquata also connects directly with Opterra for firms where operational and investment data need to reflect a consistent, validated source of truth.
Aquata follows a consumption-based pricing model, allowing firms to scale usage in line with data volumes and workflows.
When evaluating a Databricks-based build, the full cost includes DBU compute costs, cloud provider infrastructure fees, specialized engineering talent to build and maintain Spark pipelines, investment domain expertise to design the data model and quality rules, ML governance overhead, and ongoing cluster optimization, in addition to the opportunity cost of engineering time not applied to investment-differentiated work.
Firms evaluating alternatives to a Databricks-based investment data build are typically comparing how different approaches balance engineering investment against speed to operational capability.
At this stage, the focus often shifts to a small number of questions: whether a purpose-built platform can accommodate the firm's specific data requirements without requiring a multi-year engineering program, how quickly it can be deployed, and whether the combination of a governed investment data layer with downstream Databricks usage addresses both operational and quantitative requirements more efficiently than a single-platform build.
Databricks is a strong choice for AI and ML engineering workloads in investment management. The question is whether it is the right foundation for an investment operations data platform, and what it takes to build and maintain one on top of it.
This is the context in which platforms such as Aquata are typically considered.
If you are evaluating a Databricks-based investment data build, it is worth understanding how Aquata fits within a broader data architecture, and where a governed investment data layer creates more value than a custom Databricks build.
With Aquata, investment data management is handled as a maintained platform, ingestion, normalisation, quality, distribution, and analytics, without the engineering program required to build that capability on Databricks. For firms where Databricks has a role, that role is usually downstream of a governed investment data layer, not instead of one.
Arcesium works with firms at various stages of their data infrastructure evolution, including those with existing Databricks environments.
The investment required in a Databricks build does not end at go-live. Investment operations requirements change: new strategies require new data models, new counterparties require new integrations, regulatory changes require updated quality rules, and Databricks itself evolves. The build requires sustained engineering attention indefinitely. Managing Databricks cluster performance and query optimization is an ongoing discipline, not a one-time task. Managed upgrade cycles in a SaaS delivery model transfer that responsibility to the vendor, freeing engineering capacity for investment-differentiated work.