ABOUT SKETCHDECK.AI
SketchDeck.ai is transforming construction estimation through applied AI. Our platform automates complex parts of the estimation process, helping construction professionals work faster, improve accuracy, and bid on more opportunities.
Our consistent record of delivering practical AI solutions has established SketchDeck.ai as a trusted technology partner within the construction industry.
ABOUT THE ROLE
We are looking for a Senior Platform Engineer to lead the evolution of the infrastructure, deployment systems, and operational foundations behind our AI estimation platform.
This is a hands-on senior engineering role with significant ownership. You will help design and implement a secure, scalable, multi-tenant platform, modernize asynchronous job processing, automate delivery, strengthen observability, and improve the reliability of production systems across AWS and GCP.
You will work directly with the Head of Engineering and collaborate closely with backend, frontend, and machine learning engineers. We are a focused team where engineers influence technical direction, own outcomes, and see their decisions reach production quickly.
This role is ideal for someone who enjoys improving critical systems through measured, reversible changes while continuing to support active customers and product development.
WHAT SUCCESS LOOKS LIKE
During your first year, you will help us:
• Build a secure and scalable multi-tenant platform with strong tenant isolation.
• Introduce managed, reliable, and observable asynchronous job processing.
• Create a repeatable deployment process with automated validation and rollback.
• Improve visibility into system health, deployments, customer impact, and engineering performance.
• Reduce operational risk through stronger access controls, workload identity, and secrets management.
• Improve cloud efficiency using measured usage and cost data.
• Strengthen incident response, recovery procedures, and production ownership.
• Establish infrastructure practices that allow the platform to scale without creating unnecessary operational overhead.
ABOUT YOU
• You think about blast radius, failure modes, recovery time, and rollback before changing a production system.
• You select technology based on business needs, operational maturity, and total cost of ownership. You can explain your decisions clearly to technical and non-technical stakeholders.
• You automate repetitive work, document important decisions, and simplify systems whenever possible.
• You are comfortable changing critical infrastructure incrementally. You design migrations with validation, observability, and a safe path back.
• You take ownership beyond implementation. You verify that changes work in production, measure their impact, and address follow-up work.
• You communicate clearly and collaborate effectively across application, infrastructure, security, and machine learning concerns.
KEY RESPONSIBILITIES
Platform Architecture and Multi-Tenancy
• Lead the development of shared multi-tenant infrastructure, including tenant isolation, authorization boundaries, storage design, migration tooling, and progressive rollout.
• Design migration strategies that include compatibility periods, integrity checks, staged cutovers, feature controls, and tested rollback procedures.
• Partner with application engineers to ensure tenant context remains explicit and enforced across database, storage, queue, and service boundaries.
• Define architectural standards that balance security, reliability, delivery speed, and operational simplicity.
Asynchronous Processing
• Modernize asynchronous job processing using managed queue services such as AWS SQS.
• Design priority handling, per-tenant fairness, retries, visibility timeouts, dead-letter queues, and failure recovery.
• Ensure workers support idempotent processing, duplicate delivery, safe retries, and observable job states.
• Create clear operational procedures for failed, delayed, or corrupted jobs.
Deployment and Delivery
• Own the evolution of our deployment process from artifact creation through production release.
• Automate environment provisioning, configuration, deployment validation, smoke testing, rollback, and release tracking.
• Make deployments frequent, repeatable, observable, and safe.
• Improve CI/CD controls across linting, type checking, automated tests, security scanning, artifact management, and deployment approval.
• Reduce manual release steps while preserving the controls required for production systems.
Observability and Reliability
• Build an observability strategy covering structured logging, metrics, distributed tracing, dashboards, alerting, and service-level indicators.
• Implement monitoring that reflects customer impact and actionable system conditions.
• Improve deployment tracking and engineering metrics, including deployment frequency, lead time for changes, change failure rate, and recovery time.
• Establish reliability targets and help teams identify the work required to meet them.
• Partner with the ML team to strengthen the availability, monitoring, redundancy, and recovery of GPU inference workloads.
Cloud Architecture and Cost Management
• Own platform architecture across our primary AWS environment and selected GCP workloads.
• Evaluate cloud services using performance, reliability, security, operating effort, and total cost.
• Use billing and utilization data to identify waste, forecast spending, and support architecture decisions.
• Guide consolidation and migration decisions using measured costs, risks, and expected business value.
• Improve resource scheduling, rightsizing, lifecycle management, and ownership visibility.
Security and Access Management
• Modernize secrets and access management by replacing long-lived credentials with workload identity, role-based access, and short-lived tokens.
• Apply least-privilege access across services, environments, engineers, and automated systems.
• Improve auditability for infrastructure changes, production access, and deployment activity.
• Partner with engineering leadership to strengthen security controls without creating unnecessary delivery friction.
Infrastructure as Code
• Establish consistent infrastructure-as-code practices for provisioning, configuration, and environment management.
• Improve our current Ansible-based workflows while guiding the adoption of declarative infrastructure tooling where appropriate.
• Create reusable modules, documented standards, review controls, and drift detection.
• Ensure infrastructure changes follow the same review, testing, and release discipline as application code.
Incident Response and Operational Excellence
• Participate in incident response and help develop a sustainable on-call practice.
• Create playbooks for common production failures, data recovery, deployment rollback, and service restoration.
• Facilitate blameless incident reviews focused on system improvements and corrective actions.
• Ensure incident follow-up work receives clear ownership and reaches completion.
• Improve backup, restore, disaster recovery, and business continuity procedures through regular testing.
MUST-HAVE QUALIFICATIONS
• 6 or more years of experience in platform engineering, infrastructure, DevOps, SRE, or a related discipline.
• Substantial experience owning production systems and responding to operational incidents.
• Deep AWS experience across compute, networking, IAM, S3, container infrastructure, and managed services.
• The ability to design IAM policies, troubleshoot access issues, and evaluate the cost implications of architecture choices.
• Strong Linux, networking, and Docker fundamentals.
• Experience operating multi-service container environments across development, testing, staging, and production.
• Production experience with a managed message queue such as SQS, RabbitMQ, Pub/Sub, or an equivalent service.
• A strong understanding of retries, delivery guarantees, visibility timeouts, dead-letter queues, ordering, backpressure, and idempotency.
• Production experience with configuration management or infrastructure-as-code tools such as Ansible, Terraform, CloudFormation, or Pulumi.
• Experience owning CI/CD pipelines, including build automation, test gates, artifact management, deployment, validation, and rollback.
• Practical experience operating production databases, including backup, restore, replication, migration, and recovery procedures.
• Proficiency with Python sufficient to read, debug, test, and modify application and infrastructure code.
• Demonstrated experience migrating or re-architecting a live production system while protecting customer availability and data integrity.
• Strong written communication skills, including experience creating design documents, operational playbooks, and architecture decision records.
PREFERRED QUALIFICATIONS
• Experience introducing multi-tenancy into an established production platform.
• Experience designing tenant isolation across application, database, object storage, queue, and identity boundaries.
• Experience with OpenTelemetry, Prometheus, Grafana, Sentry, or comparable observability platforms.
• Experience implementing workload identity federation or OIDC-based cloud access.
• Experience replacing static cloud credentials with short-lived, role-based access.
• Experience operating infrastructure across AWS and GCP.
• Experience evaluating cloud consolidation, workload placement, data transfer, and egress costs.
• Experience operating GPU infrastructure or machine learning inference workloads.
• Experience with Nginx, reverse proxies, load balancing, TLS, and production networking.
• Experience with cloud cost optimization and measurable FinOps outcomes.
• Experience designing backup, recovery, and disaster recovery processes.
• Experience in a growing software company where you owned infrastructure without a large platform team.
• Regular use of AI-assisted development tools to improve engineering quality and delivery speed.
OUR TECHNOLOGY
Our platform includes:
• Python and Flask
• React 19 and TypeScript
• MongoDB and Amazon DocumentDB
• Redis
• Docker and Docker Compose
• AWS services, including EC2, S3, EFS, ECR, IAM, and managed infrastructure services
• Google Cloud Run and GPU compute workloads
• Bitbucket Pipelines
• Ansible
• Nginx
• Auth0
• Prometheus and Sentry
• Python-based computer vision services using YOLO-family models
Our environment includes established systems and modern services. You will help us evolve the platform thoughtfully while maintaining reliability for our customers.
WHY SKETCHDECK.AI?
• Lead platform work with clear business value and visible customer impact.
• Influence architecture, tooling, reliability, security, and cloud strategy.
• Work directly with engineering leadership and experienced product engineers.
• Take ownership from technical design through production operation.
• See meaningful architecture decisions reach production within the same quarter.
• Solve real scaling and reliability challenges in an applied AI product.
• Join a team that values autonomy, accountability, practical engineering, and clear communication.
• Help define the engineering standards and platform foundations that support our next stage of growth.
Ready to build the platform behind the future of construction estimation?
Apply at careers@sketchdeck.ai
