Jose Andrade

Jose Andrade

Technical Leader & Cloud / AI Infrastructure Architect
AI Infrastructure • Site Reliability Engineering • Observability • Enterprise Cloud Architecture

Professional Summary

Technical Leader and Cloud/AI Infrastructure Architect with 25+ years of progressive leadership directing enterprise-scale systems architecture, site reliability engineering, and data center transformations. Proven track record establishing global SRE and observability practices at Google Cloud, advising Fortune 100 executives, and leading multi-billion-dollar cloud infrastructure initiatives.

Expert in agentic observability, distributed AI operations, hybrid-cloud automation, and resilient software delivery. Bilingual public speaker, co-author of O'Reilly's Site Reliability Engineering (2nd Edition), and technical leader adept at aligning cross-functional engineering, sales, and executive teams in mission-critical initiatives.

Google Cloud Senior Staff CE 2023 President's Club Award Founder, CE SRE Practice O'Reilly Author (SRE 2nd Edition) Harvard ALM (Honors)

Core Strengths & Leadership

Proven executive, architectural, and engineering capabilities developed across Fortune 100 enterprises, hyper-growth tech, and top academic institutions:

AI Operations & SRE Leadership

Founded and scaled the global Customer Engineering Site Reliability Engineering (CE SRE) practice at Google Cloud. Formulated operational runbooks, agentic observability strategies, and SLO-based operational standards for high-throughput AI/ML model serving.

Multi-Billion-Dollar Cloud Architectures

Served as Principal Infrastructure Lead governing strategic multi-year cloud commitments valued in billions of dollars. Architected migration paths, enterprise landing zones, and blueprints for large-scale data center modernization across GCP, AWS, and Azure.

Thought Leadership & Authoring

Contributing author to O'Reilly Media's definitive Site Reliability Engineering (2nd Edition). Keynote and technical speaker at international conferences, podcast host, and founder of tech media properties read by millions worldwide.

Executive & Cross-Functional Alignment

Adept at translating deep distributed systems and reliability concepts into actionable business strategy for C-level executives, sales teams, and engineering organizations alike. Bilingual communicator (fluent in English and Spanish).


Professional Experience


Senior Staff Customer Engineer & Global Practice Lead – AI Infrastructure & SRE

Google Cloud — Go-To-Market (GTM)
Global Observability SME Founder, CE SRE Practice AI SRE Technical Program Lead 2023 President’s Club Award
2020 – Present
  • Spearheaded Google Cloud’s global observability and agentic observability strategy, designing unified reference architectures, demonstration platforms, and technical collateral adopted by field teams globally.
  • Founded and scaled the global Customer Engineering Site Reliability Engineering (CE SRE) practice, establishing technical frameworks and operational runbooks to guide strategic enterprise accounts in adopting production-grade SRE methodologies.
  • Served as Principal Infrastructure Lead on strategic multi-year cloud commitments valued in billions of dollars, architecting scalable, resilient migration paths and landing zones for enterprise workloads.
  • Formulated technical blueprints and architectural roadmaps for large-scale data center modernization and organizational cloud transformation initiatives.
  • Partnered directly with enterprise AI operations teams to audit, architect, and optimize AI/ML production pipelines, establishing SLO-based operational standards for high-throughput model serving.
  • Promoted to the highest-level engineering tier within the Specialists organization; honored with the prestigious 2023 President’s Club award in recognition of outstanding technical impact and customer success.

Senior Infrastructure Cloud Consultant & Migration Practice Lead – Americas

Google Cloud — Professional Services Organization (PSO)
Technical Program Manager Global SRE Delivery Lead
2018 – 2020
  • Directed cross-functional engineering and consulting teams executing complex cloud infrastructure foundations, application migrations, and cloud-native application architectures for Fortune 100 enterprises.
  • Appointed Migration Practice Regional Practice Lead for the Americas, governing technical delivery standards, methodology development, and engagement staffing.
  • Architected and delivered customer-facing SRE transformation programs, translating site reliability theory into actionable operational playbooks and monitoring infrastructure.
  • Scoped multi-phase enterprise consulting engagements, authoring comprehensive Statements of Work (SOWs) and managing resource allocation across complex workstreams.
  • Represented Google Cloud as a keynote and technical speaker at industry and corporate conferences, presenting on advanced infrastructure and SRE practices.

Senior Solutions Architect, DevOps Lead & Software/Systems Engineer

Yale University — ITS, Infrastructure & Design Services
2015 – 2018
  • Architected and developed "Spinup" and "SpinupManaged"—enterprise self-service hybrid cloud automation platforms (Go, Ruby, JavaScript, AWS ECS, Elasticsearch) that automated provisioning of VMs, containers, storage, and databases across AWS, Azure, VMware, and OpenStack, slashing provisioning turnaround times while incorporating automated financial chargeback tracking.
  • Founded and led the inaugural DevOps engineering team at Yale IT, successfully implementing Agile/Scrum delivery models across central infrastructure.
  • Formulated Yale's initial AWS and Azure cloud adoption roadmap, establishing administrative hierarchies, security controls, and enterprise connectivity (redundant AWS Direct Connects).
  • Served as principal cloud authority across the university, advising academic and administrative leadership on workload migrations utilizing EC2, ECS, RDS, DynamoDB, Lambda, SQS, SNS, and SES.
  • Engineered compliant, high-security computing environments meeting strict HIPAA and FERPA regulatory standards for clinical, research, and institutional data.

Systems Programmer II & Full-Stack Developer

Yale University — Department of Neuroscience
2012 – 2015
  • Designed, developed, and maintained full-stack web platforms and research portal applications (Laravel, Django, Drupal) supporting high-impact neuroscientific research initiatives.
  • Founded the Yale School of Medicine IT Partner group to foster cross-departmental technical collaboration and standardized tooling.
  • Managed heterogeneous infrastructure including multi-distribution Linux and Windows server farms supporting laboratory data workflows, endpoint management, and MFA rollouts.
  • Acted as chief technology advisor to department chairs, principal investigators, and laboratory researchers.

Software Engineer & Systems Specialist

US District Courts — District of Connecticut
2009 – 2012
  • Engineered "Dashboard," an application framework integrating with federal court docket systems (Perl, PHP, Bash, JavaScript, MySQL, Informix on RHEL/CentOS), producing custom judicial reporting and administrative workflow tools.
  • Maintained core operating systems, high-availability storage, directory services, and network infrastructure supporting federal judicial operations.
Early Career Summary
  • CompuCom Systems | Senior Lead Engineer (2004 – 2009): Managed multi-million-dollar nationwide infrastructure deployments, including completing a multi-year national server and workstation modernization program for TD Bank ahead of schedule and under budget; developed the company's first nationwide engineering collaboration intranet.
  • Andraos Capital Management / Guardian Insurance | Network & Systems Administrator (2003 – 2004): Directed multi-branch routing, switching, Active Directory, and Linux systems; rolled out early enterprise VoIP solutions to significantly reduce operational overhead.

Publications, Leadership & Technical Media

Authoring industry-standard technical literature, scaling global tech journalism, and organizing engineering communities:

Site Reliability Engineering, 2nd Edition
O'Reilly Media (2026)

Contributing Author • O'Reilly Learning

Co-authored three chapters on the Value of Reliability, Observability & Monitoring, and SLO creation for the definitive industry standard publication.

Engadget / AOL / Verizon
2005 – 2018

Founding Editor & Editor-in-Chief (Spanish Edition)

Founded and scaled the Spanish edition of Engadget (es.engadget.com) from inception to millions of monthly unique visitors; directed global editorial teams covering consumer electronics, enterprise computing, and emerging tech. Hosted a top-ranked technology podcast with thousands of downloads a week.

DevOpsCT & DevOpsDays Hartford
Regional Community

Co-Founder & Lead Organizer

Organized regional technical conferences and meetups connecting hundreds of software engineers, cloud architects, and operations practitioners across Connecticut and the Northeast.


Core Technical Competencies

Deep architectural mastery, operational experience, and software engineering capabilities:

Cloud & Distributed Infrastructure

Enterprise landing zones, multi-cloud platforms, and container orchestration:

Google Cloud (GCP) Amazon Web Services (AWS) Microsoft Azure Hybrid Architectures Kubernetes Docker & Containers Data Center Migration Infrastructure as Code (IaC) Terraform VMware OpenStack
AI Operations & Reliability (AIOps / SRE)

Production-grade reliability frameworks and AI workload governance:

Site Reliability Engineering (SRE) Agentic Observability SLI / SLO / SLA Governance AI Model Serving Reliability Incident Management Post-Mortem Facilitation TTD / TTR Reduction Toil Elimination Operational Runbooks
Observability & Telemetry

Unified telemetry, distributed systems instrumentation, and analytics:

OpenTelemetry (OTel) Prometheus Grafana Google Cloud Observability Distributed Tracing Log & Metric Analytics Dynamic Instrumentation Telemetry Correlations
Architecture, Security & Governance

Mission-critical enterprise governance and regulatory compliance:

Enterprise Systems Architecture Agile / Scrum Leadership DevSecOps HIPAA Compliance FERPA Compliance Multi-Tenant Governance Zero Trust Security Identity & Access (IAM)
Languages, Frameworks & Systems

Polyglot development and operating system internals:

Go Python Bash JavaScript / Node.js Ruby PHP Perl Java C / C++ Linux Kernel / Internals (RHEL, Debian, Ubuntu) macOS & Windows Internals RESTful APIs & Microservices SQL & NoSQL Databases

Education & Advanced Certifications

Education

Harvard University
Master of Liberal Arts (ALM) in Information Management Systems
Honors
Liberty University
Bachelor of Science in Management Information Systems
Summa Cum Laude
Los Angeles Pierce College
Associate in Science in Computer Science - Networking Technologies

Advanced Certifications

  • Google Cloud Certified: Professional Cloud Architect
  • Google Cloud Certified: Professional Cloud DevOps Engineer
  • Google Cloud Certified: Professional Cloud Developer
  • Google Cloud Certified: Professional Cloud Network Engineer
  • Cybersecurity Graduate Certificate — Harvard University
  • Web Technologies Graduate Certificate — Harvard University
  • ITIL Foundation • CompTIA A+ • Networking Technologies • Web Design

Get In Touch

Interested in discussing AI infrastructure, site reliability engineering, cloud architecture, or advisory opportunities? Feel free to send a message below.


About This Site

This site is built serverless and hosted on a Google Cloud Storage (GCS) bucket with high-availability TLS delivery via Google Cloud Global External Load Balancing. The dynamic contact form is powered asynchronously by a serverless backend on AWS API Gateway & AWS Lambda.