Data Engineer @Azumo
Data and Analytics
Salary unspecified
Remote Location
Employment Type full-time
Posted 2wks ago

[Hiring] Data Engineer @Azumo

2wks ago - Azumo is hiring a remote Data Engineer. πŸ’Έ Salary: unspecified πŸ“Location: Latin America (LATAM)

Role Description

Azumo is hiring a Data Engineer to own the layer everything else depends on: ingestion and transformation pipelines, storage and warehouse design, and the retrieval infrastructure that AI systems query. The role is fully remote across Latin America, aligned to your client's working day.

You will not be building pipelines that work until the schema changes. Azumo has shipped production data systems since 2016, and the work here is judged downstream:

  • Whether a model can be trained on what you deliver.
  • Whether a retrieval query returns the right passage.
  • Whether a number survives being questioned by the client.

This role sits in Azumo's engineering organization, which is built around four lanes:

  • Data Scientist lane: owns the question and the method.
  • AI Engineer lane: owns production behavior.
  • Software Engineer lane: owns AI-augmented product delivery.
  • Data Engineer lane: owns pipelines, storage, and the retrieval layer.

One question places the boundary: when the output is wrong, whose problem is it?

  • "The data was missing, stale, or wrong by the time it arrived" is yours.
  • "The system did the wrong thing with data that was correct" is the AI Engineer's.

What you will build:

  • Ingestion and transformation pipelines: Batch and streaming ingestion on Spark, Kafka, dbt, and Airflow, with idempotency, backfills, schema evolution, and late-arriving data handled by design rather than by hand.
  • Storage and modeling: Warehouse and lakehouse design on Snowflake, BigQuery, Redshift, or Databricks, with partitioning, file layout, and query cost treated as engineering decisions.
  • The retrieval layer: The chunking, embedding, and indexing pipelines that feed RAG systems on pgvector, Pinecone, Qdrant, or Azure AI Search, and the freshness, deduplication, and permission problems that come with them.
  • Data quality as a contract: Tests, expectations, lineage, and alerting. If a pipeline is wrong, the people downstream should hear it from you and not from the client.
  • Sensitive data by default: PII classification, masking, row and column level access, retention and deletion, and an audit trail that holds up when a client asks who read what.
  • Production operation: Containerized deployment on Azure or AWS, CI/CD, orchestration, observability, and explicit cost and runtime budgets that you own rather than discover after the invoice.
  • Work inside the client's environment: Their repositories, their standups, sometimes their customer calls. Azumo is SOC 2 certified, client code stays in client repositories, and some engagements carry additional requirements such as HIPAA.

Our engineers build with AI every day. Claude Code, Codex, and similar tools are part of the standard toolchain here, not an experiment. We run an automated audit across the whole codebase on day one and every day after, grading security, cost, and architecture findings by severity with the exact file and line, so a small team can move quickly without quality drifting. We stay vendor-neutral across OpenAI, Anthropic, and open-weight models, and we run Valkyrie, our own production layer, when a single interface to any model is the right call.

Qualifications

  • 5+ years building and operating production data pipelines, with Python and SQL as your primary languages, plus the engineering fundamentals that go with it: testing, code review, CI/CD, Git, containers, and orchestration.
  • Deep expertise in designing and building data warehouses or lakehouses, including dimensional modeling, incremental processing, and the cost and performance trade-offs behind each choice.
  • Distributed processing at production scale with Spark, Kafka, Flink, or equivalent, including the failure modes that only appear under load.
  • Orchestration as an engineering discipline rather than a cron replacement: Airflow, Dagster, or Prefect, with retries, idempotency, and backfill strategy you can defend.
  • Transformation under version control, with tests and lineage: dbt or something you built yourself.
  • Cloud deployment experience, Azure preferred and AWS acceptable, with Docker, CI/CD pipelines, and infrastructure as code (GitHub Actions, Terraform, or Bicep).
  • Working discipline around pipeline cost, runtime, and throughput. You can explain what a pipeline costs to run and what you did about it.
  • Active use of AI-assisted coding tools such as Claude Code, Cursor, or GitHub Copilot in real delivery work.
  • Clear written and spoken English, C1 or above, and the confidence to explain a technical trade-off directly to a client.
  • Bachelor's degree in Computer Science, Data Science, or a related field, or equivalent professional experience.

Preferred Qualifications

  • Vector and retrieval infrastructure: pgvector, Pinecone, Qdrant, FAISS, or Azure AI Search, and the retrieval-quality problems that come with it.
  • Streaming, real-time, or high-throughput workloads.
  • Experience with cloud-based managed services like Airflow, Glue, Elastic stack, Amazon Redshift, Snowflake, BigQuery, Azure SQL Db, EMR, Databricks.
  • Prior experience with notebooks using Jupyter, Google Collab, or similar.
  • Delivery under a compliance regime such as SOC 2 or HIPAA.
  • Contributions to open-source data libraries, published technical writing, or active participation in the data engineering community.

Benefits

  • Paid time off (PTO)
  • U.S. Holidays
  • AI Training
  • Mentored career development
  • Profit sharing
  • $US remuneration
Before You Apply
️
remote Be aware of the location restriction for this remote position: Latin America (LATAM)
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Data Engineer @Azumo
Data and Analytics
Salary unspecified
Remote Location
Employment Type full-time
Posted 2wks ago
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
️
remote Be aware of the location restriction for this remote position: Latin America (LATAM)
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
Γ—
Apply to the best remote jobs
before everyone else

Access 125,000+ vetted remote jobs and get daily alerts.

4.9 β˜…β˜…β˜…β˜…β˜… from 500+ reviews

⚑ 127,118+ remote jobs, refreshed hourly

πŸ”” Real-time alerts: Apply first, direct to employer

πŸ›‘οΈ Vetted companies, no scams, true remote only

Unlock All Jobs Now

Maybe later