T
ThredUp
DevopsSenior
Senior Engineer, Infrastructure
Apache AirflowArgoCDAmazon AuroraAutovacuumAWSBashAWS CloudWatchCloudflareDatabricksDatadogDNSDynamoDBEXPLAINEXPLAIN ANALYZEgh-ostGitHub ActionsIAMIAM DB authenticationIncident ResponseIstioJenkinsKafkaKubernetesMongoDBMySQLNetwork Securityonline DDLpartition-based retentionpg_repackPostgreSQLpt-online-schema-changePythonRabbitMQAmazon RDSRollbarService Meshsnapshot restoresTeleportTerragruntTerraformTest AutomationVPCVPNVulnerability Management
About the Position
ThredUp is a large online resale platform. We are seeking a Senior Infrastructure Engineer to design, build, and evolve core infrastructure with a focus on database reliability, using AWS, Kubernetes, Terraform, and automation. The role involves leading cross-team initiatives and ensuring systems are scalable, cost-effective, and secure.
Responsibilities
- Lead or significantly contribute to medium-to-large infrastructure projects crossing multiple engineering teams
- Serve as a domain expert in cloud infrastructure, orchestration, observability, and platform automation, and as the team’s primary owner of database reliability and performance
- Ensure ThredUP’s infrastructure evolves to support scale, resilience, and developer productivity
- Architect and implement highly available, secure, and cost-efficient cloud infrastructure using AWS, Kubernetes (EKS), and Terraform
- Drive improvements in CI/CD, observability, networking, and security across the platform
- Provide high-quality, impactful technical contributions across infrastructure projects, setting engineering standards
- Participate in and lead design reviews, providing constructive feedback and driving engineering excellence
- Operate, monitor, and tune AWS Aurora and RDS clusters (MySQL and PostgreSQL), including parameter groups, maintenance, minor/major upgrades, and point-in-time restores
- Own HA and replication behavior: respond to failovers, work with cluster vs. instance endpoints, and run the checks required before promoting a reader to writer
- Triage and resolve CPU/IO/locking/replication incidents; analyze slow query logs and EXPLAIN/EXPLAIN ANALYZE plans to produce immediate mitigations and long-term fixes
- Plan and execute zero-downtime schema changes with safe rollback paths (online DDL, gh-ost / pt-online-schema-change, pg_repack, logical replication, trigger-based backfills)
- Partner with application teams to remove N+1 queries, tune indexes, and rewrite inefficient predicates; design partitioning and retention for high-write tables
- Design, execute, and regularly validate backup and disaster-recovery procedures
- Act as an expert in infrastructure design, performance, and operations across multiple systems and services
- Promote shared ownership of infrastructure by driving documentation, tooling, and process improvements
- Monitor and optimize system performance, ensuring reliability while protecting teams from burnout
- Build relationships with engineering teams, product managers, and cross-functional partners to ensure infrastructure supports company goals
- Contribute to defining strategic technical direction, setting infrastructure roadmaps, and guiding prioritization
- Advocate for best practices in reliability, security, and cost management
- Ensure knowledge is shared within the team, reducing single points of failure
Requirements
- 6+ years of relevant industry experience with a Bachelor’s degree in Computer Science, Engineering, or related field (or equivalent experience)
- Hands-on experience operating MySQL and PostgreSQL in production (Aurora/RDS preferred), including indexing strategies, transactions, locking/MVCC, and performance tuning
- Demonstrated experience performing zero-downtime schema changes with safe rollbacks
- Proven track record of designing and scaling infrastructure for distributed, service-oriented architectures
- Expertise in AWS (EKS, RDS, IAM, cost optimization)
- Proficiency with Kubernetes and Terraform
- Experience with CI/CD pipelines and practices (e.g., Jenkins, GitHub Actions, or ArgoCD – one or more)
- Experience with observability and monitoring tooling (Datadog, CloudWatch, or similar), including slow query and error log analysis
- Ability to diagnose connectivity issues affecting database access (security groups, VPC routes/TGW, VPN, DNS)
- Scripting and automation skills (Python, Bash, or similar)
- Excellent communication and collaboration skills
- Additional plus: Experience with Teleport, IAM DB authentication, tuning Postgres autovacuum, designing partition-based retention policies, Terraform/Terragrunt module design for RDS/Aurora clusters, service mesh (Istio), edge/WAF tooling (Cloudflare), data infrastructure (Airflow, Databricks), messaging/streaming systems (RabbitMQ, Kafka, DynamoDB, MongoDB), network security, vulnerability management, incident response, Rollbar, centralized logging pipelines, fintech, testing automation, compliance-heavy environments (GDPR, SOC2), prior leadership in scaling platforms for e-commerce or high-growth startups
Benefits
- Monthly allowance for insurance/education ($200 gross)
- 50% paid sabbatical after 3 year anniversary
- Paid parental leave for new mothers and fathers
- IT Kit
Senior Engineer, Infrastructure
View Original