• About us
    About us
    question mark
    Who we are

    Learn more about Mantu values, governance and offices.

    hexagon
    Our brands

    11 brands united by a shared vision.

    leave
    Sustainability

    Our strategy through diversity, environment and innovation.

    bookshelf
    Pressroom

    Breakthroughs, partnerships, and voices behind the transformation.

  • What we do
    What we do
    mantu
    PRACTICES

    Four practices designed to empower organizations, connect talent, and shape sustainable growth.

    cpu
    Technology

    Deep industry knowledge & cutting edge technology to co-create meaningful solutions.

    handshake
    Total Talent Management

    Tech to boost talent and create strong links between companies and the minds they need.

    digital qr
    Creative Intelligence

    Ensure continuity between decision, activation, and adoption. One team, one trajectory, through to lasting impact.

    medal
    Leadership & Advocacy

    Equip executive teams to define their purpose, shape their positioning and drive their strategy.

  • Insights
    Insights
    book open 4
    Blog

    Bold thinking. Fresh perspectives.

    book check
    Client Stories

    Where audacious ideas turn into real stories.

    mantu best managed companies award
    Mantu awarded one of Switzerland’s Best Managed Companies 2025 by Deloitte

    This award highlights the exceptional performance of privately held Swiss companies that demonstrate excellence in strategy, governance, innovation, and long-term results.

    Read more
    WeMeet 2025-2772 1 1
    Mantu signs the DEI Charter

    At the beginning of July 2025, Mantu’s Executive Committee signed the DEI Charter to foster diversity, equity, and inclusion at Mantu.

    Read more
  • Careers
    Careers
    binoculars
    Life at Mantu

    Mantu, as seen by its team members.

    building
    Find a company

    Mantu brings together complementary brands that cover many sectors, all around the world.

What-Is-Data-Orchestration-

What Is Data Orchestration?

Data orchestration is the automated coordination of data pipelines: scheduling, dependencies, and error handling, without manual intervention.

Without it, teams manage these dependencies manually. Wakefield Research found that data engineers spend 44% of their time on pipeline maintenance, at a cost of roughly $520,000 a year per team.

Why data pipelines outgrow a cron job


Every data pipeline starts simple: extract, transform, load, on a schedule. A cron job handles that fine, until the pipeline has a dependency. The transform step can't run until the extract step finishes. The load step should only run if the transform step produced valid output. Add a second pipeline that shares a source table, a retry policy for the API that occasionally times out, and a need to know exactly which job failed at 3am and why, and cron stops being an answer.

An orchestration layer exists to manage exactly that complexity. It doesn't just run jobs, it understands the relationships between them, expressed usually as a directed acyclic graph (DAG), and it decides what runs next based on what already happened, not just what time it is.

Orchestration vs workflow automation: not the same problem


Data orchestration gets lumped in with generic workflow automation tools, and that's a mistake worth correcting early. A workflow automation tool like a business process platform routes an approval or moves a file between apps, triggered by a human action or a simple event. Data orchestration coordinates programmatic jobs across data infrastructure: Spark jobs, dbt models, warehouse loads, ML feature pipelines, often processing millions of rows with strict ordering requirements and no human in the loop.

The distinction matters for the buying decision. A business team automating approvals doesn't need DAG-based dependency resolution or data lineage tracking. A data engineering team running 200 interdependent pipelines needs exactly that, and a generic automation tool won't give it to them.

What a modern orchestration layer actually does


A capable orchestration platform, whether that's Apache Airflow, Dagster, Prefect, or a managed service like Azure Data Factory, covers five capabilities that separate it from a scheduler:

  • Dependency management. Jobs run in the correct order based on their actual data dependencies, not a fixed clock time that assumes the upstream job always finishes on schedule.

  • Retry and failure handling. A transient API timeout triggers an automatic retry with backoff, while a genuine data quality failure halts the pipeline before bad data reaches a downstream table.

  • Observability and lineage. Every run is logged with enough detail to answer "which pipeline touched this table, and when" without someone reconstructing it from memory.

  • Event-driven triggers. Pipelines can start based on a file landing in storage or an upstream event firing, not just a fixed time window, which matters once data arrives at unpredictable intervals.

  • Resource and concurrency control. Multiple pipelines competing for the same compute or the same source system get scheduled to avoid overload, rather than all firing at once and taking the source database down.

Pipeline reliability comes from these five working together. A tool that handles scheduling well but has no real lineage tracking still leaves a team debugging incidents by grepping through logs.

Data orchestration and the shift toward data mesh


Data mesh enablement puts new pressure on the orchestration layer. In a mesh model, data ownership moves to domain teams, marketing, finance, product, each running their own pipelines instead of funneling everything through one central data engineering team. Cross-domain integration only works if pipelines owned by different teams can depend on each other's outputs without someone manually coordinating handoffs over Slack.

This is where data flow governance and orchestration start to overlap. A domain team publishing a dataset needs the orchestration layer to enforce contracts, this pipeline expects that schema, this SLA, so that a downstream consumer's job doesn't silently break when the upstream team changes a column name. Event-driven architecture fits naturally here: a domain publishes a "dataset updated" event, and every downstream pipeline that depends on it triggers automatically, instead of running on a fixed schedule and hoping the upstream data landed in time.

Organizations moving from a centralized warehouse to a domain-oriented model often underestimate how much of that transition depends on getting orchestration right first. Mantu's data strategy consulting practice works through this sequencing with clients, since a mesh rollout without a solid orchestration foundation tends to just distribute the pipeline chaos instead of fixing it.

Choosing between orchestration tools


The tool decision usually comes down to three questions: how the team wants to define pipelines (Python code versus a visual interface), how much operational overhead the team can absorb (self-hosted Airflow needs infrastructure management; a managed service trades some flexibility for less maintenance), and how deeply the tool needs to integrate with the existing stack.

Tool type

Best fit

Trade-off

Code-first (Airflow, Dagster, Prefect)

Teams with strong engineering resources, complex dependency graphs

Requires infrastructure to run and maintain

Managed cloud service (Azure Data Factory, AWS Step Functions)

Teams already committed to one cloud provider

Less flexibility outside that ecosystem

Low-code integration platforms

Smaller teams, simpler pipelines, faster setup

Struggles at scale with complex dependencies

None of these choices is permanent, and none of them fixes a pipeline architecture that was never designed with dependencies and failure handling in mind to begin with. The tool automates the coordination. It doesn't decide what should be coordinated with what.

Getting the foundation right before scaling pipelines


Teams that bolt orchestration onto an existing mess of scripts usually end up automating the chaos rather than removing it. The pipelines that benefit most from an orchestration layer are the ones designed with clear ownership, documented dependencies, and defined data contracts from the start.

Getting that foundation in place, before choosing a specific orchestration tool, is typically where Mantu's data strategy and operating model experts come in: mapping which pipelines actually depend on each other, who owns each one, and what a broken dependency should trigger, so the orchestration layer has something coherent to coordinate in the first place.