Data Scientist @Rice University
Data and Analytics
Salary up to $105,000 ..
Remote Location
🇺🇸 USA Only
Employment Type full-time
Posted YDay

[Hiring] Data Scientist @Rice University

YDay - Rice University is hiring a remote Data Scientist. 💸 Salary: up to $105,000 annually 📍Location: USA

Role Description

The Rice University Department of Computer Science and Department of Materials Science and NanoEngineering are seeking a Data Scientist to build and operate the data infrastructure for READINESS, a new $20 million NSF-funded project transforming four materials synthesis systems (CVD, thermal, PVD, and plasma CVD reactors for 2D materials, oxides, and diamond films) into an AI-driven, remotely accessible autonomous laboratory.

The READINESS facility will generate substantial and diverse data, including:

  • Growth recipes
  • Reactor time-series
  • User interactions
  • Optical, SEM/TEM and AFM images
  • Spectra
  • Continuous robot telemetry

The Data Scientist will be responsible for:

  • Developing and maintaining the central data lake that captures, indexes, secures, and serves this information across the project.
  • Combining data engineering and analysis, including deploying an S3-compatible object store.
  • Developing ingestion connectors for live instruments and robot controllers.
  • Establishing SQL and vector-search capabilities under Rice single sign-on.
  • Computing embeddings over images and spectra.
  • Developing retrieval and similarity-search services.
  • Building AI agents capable of answering researchers' questions directly from the data.

Reporting to the faculty leads of the READINESS Data Infrastructure working group, the Data Scientist will collaborate closely with:

  • Synthesis platform teams
  • Robotics and AI/digital-twin working groups
  • Rice IT and Information Security
  • Graduate and undergraduate students involved in the project

This is a hands-on computer science and data engineering position that requires substantial ownership of the project's central data infrastructure.

A background in materials science, chemistry, or laboratory science is not required; domain-specific knowledge can be developed through collaboration with the project's scientific experts.

Working within a small team, the Data Scientist will develop the foundational systems through which data generated across the READINESS facility is captured, organized, and made accessible.

Qualifications

  • Bachelor’s degree in computer science, data science, engineering, or a related quantitative field.
  • Three or more (3+) years of related professional experience in data engineering, data science, software engineering, or research computing.

Requirements

  • Strong programming skills and ability to write production-quality, tested, documented code.
  • Proficiency with SQL and relational data modeling, including scientific or operational database schemas.
  • Experience with object storage (S3, Ceph, or MinIO), columnar formats (Parquet), and distributed query engines (Trino, Presto, Spark, or similar).
  • Working knowledge of Linux administration, containers (Docker/Kubernetes), and at least one major cloud platform (AWS or Azure).
  • Familiarity with PyTorch, embedding models, vector databases, and retrieval-augmented or agentic LLM applications.
  • Understanding of authentication and access controls (SSO/SAML/OIDC, role-based access, audit logging) and secure research-data handling.
  • Ability to connect software to physical instruments and devices using network or file-based protocols; ROS/rosbag experience is a plus.
  • Ability to learn an unfamiliar scientific domain, translate collaborators’ requirements into working systems, manage competing requests, and communicate clearly.

Preferences

  • Five or more years of professional experience building and operating data systems.
  • Experience building data pipelines for scientific instruments, laboratory automation, manufacturing, or IoT/telemetry.
  • Experience operating Ceph or comparable software-defined storage, or managing cloud object-storage tenancies at 100 TB+ scale.
  • Experience with GPU-accelerated similarity search (FAISS, Milvus, Qdrant) and serving open-weight LLMs.
  • Experience in academic research or national laboratories, including working with institutional IT and security offices.
  • Experience monitoring production data systems and implementing backup and recovery.

Essential Functions

  • Designs, deploys, and operates the READINESS central data lake, including the S3-compatible object store (self-hosted Ceph or cloud), bucket layout, versioning, and access policies.
  • Develops and maintains ingestion connectors that capture data unattended from synthesis tools, characterization instruments, and robotic systems, and land it in the store with structured, de-identified metadata and provenance links.
  • Designs and maintains the project's metadata and provenance schema and the Python client library that all project teams use to read and write data.
  • Deploys and administers the SQL query layer (Trino and Hive Metastore over Parquet) and programmatic APIs for researchers, the user portal, and digital-twin and AI teams.
  • Builds embedding pipelines and GPU-resident vector indexes over images, spectra, and logs, and exposes similarity-search and retrieval-augmented-generation services.
  • Develops and demonstrates AI agents that answer questions about experiments by retrieving from the data lake, in collaboration with the AI & Digital Twins working group.
  • Implements and maintains security controls, including Rice single sign-on with MFA, tiered role-based access, encryption, and immutable audit logging, and works with Rice Information Security and Research Security on design reviews and compliance.
  • Monitors system health, performs integrity checks and backups, and documents and tests recovery procedures.
  • Gathers requirements from synthesis platform, characterization, and robotics teams; attends platform team and working-group meetings; and sets node-wide standards for data formats, metadata, and access.
  • Writes documentation, user guides, and onboarding materials, and trains researchers and partner-site collaborators to use the data infrastructure.
  • Mentors graduate and undergraduate students working on data infrastructure projects.
  • Performs all other related duties as assigned.

Workplace Requirements

This position is fully remote, permitting all tasks to be completed from any location within the United States. Working hours will remain central standard time. Per Rice policy 440, work arrangements may be subject to change.

Hiring Range

Up to $105,000 annually. This is a one-year term limited, benefits-eligible position funded by a grant, soft and/or restricted funds, and may be renewed based on continued funding availability, research needs, and performance.

Before You Apply
️
🇺🇸 Be aware of the location restriction for this remote position: USA Only
‼ Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Data Scientist @Rice University
Data and Analytics
Salary up to $105,000 ..
Remote Location
🇺🇸 USA Only
Employment Type full-time
Posted YDay
Apply for this position
Did not apply ✓
Applied ✓
Sent Follow-Up ✓
Interview Scheduled ✓
Interview Completed ✓
Offer Accepted ✓
Offer Declined ✓
Application Denied ✓
Unlock 125,000+ Remote Jobs
️
🇺🇸 Be aware of the location restriction for this remote position: USA Only
‼ Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply ✓
Applied ✓
Sent Follow-Up ✓
Interview Scheduled ✓
Interview Completed ✓
Offer Accepted ✓
Offer Declined ✓
Application Denied ✓
Unlock 125,000+ Remote Jobs
×
Apply to the best remote jobs
before everyone else

Access 125,000+ vetted remote jobs and get daily alerts.

4.9 ★★★★★ from 500+ reviews

⚡ 126,168+ remote jobs, refreshed hourly

🔔 Real-time alerts: Apply first, direct to employer

🛡️ Vetted companies, no scams, true remote only

Unlock All Jobs Now

Maybe later