Role Description
Greystar is seeking a Platform Engineer to join the Data team. This is a senior, deeply technical platform operations role responsible for the reliability, governance, and cost efficiency of Greystarβs enterprise data infrastructure β a Databricks-native medallion architecture (Bronze β Silver β Gold) running entirely on Microsoft Azure. You will work within the DataOps pod under the Analytics Engineering umbrella, owning Databricks workspace administration, Unity Catalog governance, Azure infrastructure operations, and multi-system platform support.
Job Description
Databricks Platform Engineering & Administration
-
Own Databricks workspace administration: upgrades, cluster policies, compute configurations, and Delta table lifecycle management
-
Administer Unity Catalog: access controls, service principal (SP), group, and user provisioning, governance policies, and data lineage
-
Tune Spark jobs for performance, reliability, and cost β profiling bottlenecks, optimizing partitioning, managing Z-ordering, and controlling compute spend
-
Architect and enforce DMP platform standards: naming conventions, schema evolution policies, SLA tiers, and medallion layer contracts
-
Leverage Databricks Mosaic AI and Genie to build AI-native DataOps capabilities including intelligent pipeline monitoring and anomaly detection
Azure Infrastructure & Integration
-
Operate the full Azure data services stack: ADLS Gen2, Azure Data Factory (ADF), Azure Monitor, Log Analytics, Key Vault, and Event Hub
-
Collaborate with Azure infrastructure and cloud engineering teams on networking, identity, security, and resource provisioning
-
Drive cost governance through Azure Cost Management, Databricks DBU optimization, and storage lifecycle policies
Multi-Platform Support
-
Provide platform support for various systems and platforms including access management, storage optimization, and compliance
-
Support additional platforms including ADF, Synapse Analytics, Cosmos DB, and Azure SQL Server
-
Manage storage allocation, cluster optimization, and resource provisioning across all supported platforms
-
Support compliance, audit readiness, and Key Vault secret rotation and management
CI/CD & Environment Deployments
-
Own the full deployment pipeline for DMP data workflows β promoting changes from development through staging to production
-
Build and maintain CI/CD workflows using GitHub Enterprise: branch strategies, PR automation, environment-specific configuration management, and release gating
-
Enforce deployment standards: automated testing gates, rollback procedures, change documentation, and environment parity controls
-
Use Linear for sprint planning, release tracking, and issue management across deployment cycles
Identity, Access & Governance
-
Manage Azure Active Directory (AAD) service principals, RBAC assignments, and Unity Catalog permission policies
-
Partner with the Data Governance team to enforce data contracts, ownership standards, and quality SLAs within Unity Catalog
-
Maintain data lineage, metadata management, and governance framework documentation
-
Ensure platform configurations meet security, compliance, and audit requirements
Environment Parity & Dev/Staging/Prod Management
-
Own workspace configuration parity across Dev, Staging, and Production Databricks environments β cluster policies, Unity Catalog hierarchy, access controls, and compute configurations
-
Define and enforce environment promotion hygiene: configuration drift detection, environment-specific secret management, and pre-production validation gates
-
Maintain infrastructure-as-code (IaC) scripts and configuration templates to ensure reproducible, auditable environment provisioning across all tiers
Documentation & Collaboration
-
Produce thorough technical documentation: runbooks, deployment playbooks, incident post-mortems, ADRs, and platform specs
-
Partner with analytics engineers, data governance, and product stakeholders on platform design and capacity planning
-
Participate in on-call rotation and support SLA commitments for business-critical DMP data domains
AI-Driven Platform Observability & Self-Healing Infrastructure
-
Build and maintain AI-powered platform observability infrastructure β using LLMs and ML models to detect pipeline drift, classify anomalies, predict SLA risk, and generate automated incident summaries at the platform layer
-
Architect and deploy self-healing pipeline capabilities β automated retry logic, circuit breakers, and remediation triggers β as production-grade platform features, not experiments
-
Integrate Databricks Mosaic AI and Genie into operational platform workflows for intelligent cluster management, anomaly detection, and natural language platform diagnostics
-
Contribute to Greystarβs 18-month agentic AI roadmap, leading near-term delivery of platform-level self-healing and proactive alerting capabilities
Data Pipeline Architecture & Platform Standards
-
Own DMP pipeline architecture patterns from an infrastructure standpoint β Delta Live Tables (DLT) design standards, schema evolution enforcement, and medallion layer contracts
-
Define and enforce platform-wide naming conventions, partitioning strategies, Z-ordering policies, and Delta table optimization schedules (VACUUM, OPTIMIZE)
-
Partner with analytics engineers to ensure pipeline designs conform to platform infrastructure standards, SLA tiers, and compute budget constraints
Disaster Recovery & Business Continuity
-
Design and maintain DR strategies for the DMP platform β including Delta table backup policies, ADLS Gen2 geo-redundancy configuration, and Databricks workspace recovery runbooks
-
Define RTO and RPO targets for business-critical DMP data domains; test and validate recovery procedures on a recurring basis
-
Maintain business continuity plans aligned with Azure region failover capabilities and Greystarβs enterprise SLA commitments
Qualifications
-
5+ years of platform engineering, infrastructure engineering, or data operations experience in a production environment
-
Expert-level Databricks administration: Unity Catalog, cluster policies, workspace management, compute configuration, and Delta Lake internals
-
Strong command of the Azure data services ecosystem: ADLS Gen2, ADF, Azure Monitor, Log Analytics, Key Vault, Cost Management, and Event Hub
-
Identity and access management: Azure Active Directory, service principals, RBAC, and Unity Catalog governance policies
-
Hands-on experience with ADF pipeline design and orchestration at scale
-
Spark performance tuning: partitioning, Z-ordering, compute spend optimization, and bottleneck profiling
-
Proven CI/CD experience using GitHub Enterprise: environment promotion, rollback, and release management for data pipelines
-
Python and/or SQL for automation scripts, compliance validation, and platform tooling
-
Knowledge of medallion / lakehouse architecture patterns and multi-environment deployment discipline
-
Terraform / Infrastructure-as-Code (IaC): provisioning and managing Databricks workspaces, Azure resources, Unity Catalog configurations, and environment-specific infrastructure through version-controlled IaC
-
Azure DevOps Pipelines: build and release pipeline configuration, YAML pipeline authoring, and integration with GitHub Enterprise for end-to-end deployment automation
-
Azure networking: VNet configuration, private endpoints for Databricks and ADLS Gen2, NSG rules, firewall policies, and DNS resolution for secure data platform connectivity
-
Azure Monitor & Log Analytics: KQL query authoring, alert rule configuration, diagnostic settings, workbook creation, and cluster event monitoring for proactive platform health management
-
Delta Lake internals: VACUUM, OPTIMIZE, ZORDER, time travel, table statistics, and transaction log management for production-grade Delta table lifecycle governance
Requirements
-
Experience with Synapse Analytics, Cosmos DB, or Azure SQL Server in production data environments
-
Familiarity with Databricks Mosaic AI, Genie, or AI-native platform observability capabilities
-
Background in data lineage tooling, metadata management, and Unity Catalog governance frameworks
-
Experience with ERP system integrations: Yardi, Entrata, or RealPage in multi-tenant environments
-
Background in legacy BI migration or platform modernization programs
-
Linear for engineering sprint and release management
Benefits
-
Competitive Medical, Dental, Vision, and Disability & Life insurance benefits
-
Low (free basic) employee Medical costs for employee-only coverage; costs discounted after 3 and 5 years of service
-
Generous Paid Time off: 15 days of vacation, 4 personal days, 10 sick days, and 11 paid holidays
-
Birthday off after 1 year of service
-
Additional vacation accrued with tenure
-
6-Week Paid Sabbatical after 10 years of service (and every 5 years thereafter)
-
401(k) with Company Match up to 6% of pay after 6 months of service
-
Paid Parental Leave and lifetime Fertility Benefit reimbursement up to $10,000
-
Employee Assistance Program
-
Critical Illness, Accident, Hospital Indemnity, Pet Insurance and Legal Plans
-
Charitable giving program and benefits