Principal AI Cloud Storage Engineer @Blaze Talent
Artificial Intelligence
Salary unspecified
Remote Location
πŸ‡ΊπŸ‡Έ USA Only
Employment Type full-time
Posted Today

[Hiring] Principal AI Cloud Storage Engineer @Blaze Talent

Today - Blaze Talent is hiring a remote Principal AI Cloud Storage Engineer. πŸ’Έ Salary: unspecified πŸ“Location: USA

Role Description

We're looking for a Principal Software Engineer to define and build Neo Cloud's AI cloud storage platform. This is a senior individual-contributor role for an experienced system architect who can operate at the intersection of customer needs, system architecture, and production operations β€” translating what AI/ML customers will need in the future, into a storage system that is performant, resilient, and economical at massive scale.

Key Responsibilities

  • Identifying customer requirements:
    • Engage directly with customers, solutions architects, and product teams to understand future storage requirements for AI/ML workloads.
    • Extrapolate from current usage patterns and industry trends to anticipate future requirements.
    • Partner with product management to prioritize platform investments based on near-term customer needs and longer-term strategic bets.
  • System design, implementation, and operations:
    • Design, implement, and operate AI cloud object and file storage systems, including the data path, metadata path, and control plane.
    • Take a system-level approach that accounts for the full characteristics of AI workloads, building end-to-end solutions.
    • Architect for the specific demands of AI workloads: very high aggregate throughput, support for massive numbers of small and large objects and files, efficient checkpointing at scale, and predictable tail latency under heavy concurrent load.
    • Drive core storage system design decisions, including durability and consistency models, erasure coding and replication strategies, metadata scalability, multi-tenancy and isolation, and S3-compatible and POSIX/file-protocol API design.
    • Take end-to-end ownership of services in production: build for observability and operability from day one, participate in on-call, lead incident response and root-cause analysis for critical issues, and drive long-term reliability and performance improvements.
    • Identify and eliminate performance bottlenecks and scalability limits before they become customer-facing problems; lead capacity planning for rapid growth.
    • Deep understanding of data privacy and security and its implication on performance.
    • Ability to partner with network engineers to deliver complete AI storage system.
  • Technical leadership:
    • Strong bias to action and resolution of technical decisions.
    • Set technical direction and best practices for the storage organization; author and review design documents for significant architectural changes.
    • Provide deep technical mentorship to senior and staff engineers; raise the engineering bar across the team through code review, design review, and hands-on collaboration.
    • Influence technical strategy across adjacent teams.

Qualifications

  • 10+ years of professional software engineering experience building and operating cloud storage systems in production.
  • Direct experience with AI-focused storage platforms such as DDN, Weka, or VAST, including their architectural approaches to throughput, caching, and GPU-cluster integration.
  • Deep understanding of distributed systems fundamentals: consistency models, replication, consensus, partitioning, failure detection, and recovery.
  • Proven experience operating high-scale distributed systems in production, including on-call ownership, incident response, and driving systemic reliability improvements.
  • Strong systems programming skills (e.g., Go, C++, Rust, or Java) and comfort working across the stack from low-level I/O and networking to distributed control planes.
  • Excellent written and verbal communication skills.

Nice to have

  • Experience designing or tuning local node caching layers to accelerate AI training and inference data access.
  • Experience with high-performance networking.
  • Contributions to open-source storage projects, relevant patents, or published technical papers/talks.
Before You Apply
️
πŸ‡ΊπŸ‡Έ Be aware of the location restriction for this remote position: USA Only
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Principal AI Cloud Storage Engineer @Blaze Talent
Artificial Intelligence
Salary unspecified
Remote Location
πŸ‡ΊπŸ‡Έ USA Only
Employment Type full-time
Posted Today
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
️
πŸ‡ΊπŸ‡Έ Be aware of the location restriction for this remote position: USA Only
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
Γ—
Apply to the best remote jobs
before everyone else

Access 125,000+ vetted remote jobs and get daily alerts.

4.9 β˜…β˜…β˜…β˜…β˜… from 500+ reviews

⚑ 129,357+ remote jobs, refreshed hourly

πŸ”” Real-time alerts: Apply first, direct to employer

πŸ›‘οΈ Vetted companies, no scams, true remote only

Unlock All Jobs Now

Maybe later