Skip to content

Work

I turn repeated problems into software, software into platforms, and platforms into teams.

TELUS

Distinguished Engineer

2021–PresentNelson, BC

I started at TELUS by bringing more than 4,000 engineers onto a common developer experience. A weekend prototype then became Fuel iX; today I'm building sovereign compute and shared AI infrastructure across regulated institutions. I build the first path in code, then help teams carry it at enterprise scale.

Selected workView

AI Consortium

2026–Present

I lead TELUS's engineering contribution to the Agentic Control Plane, shared infrastructure for governing production AI across the Consortium.

Tokens per month
2T+
Founding members
4

Lightworks, Scotiabank, Sun Life, and TELUS build and govern the Agentic Control Plane together, and every member holds perpetual rights to what we create. I lead TELUS's engineering contribution.

The platform is the control layer for agents, models, users, and inference pipelines: approval workflows, policy enforcement, machine identity, and audit trails that let a bank or an insurer put agents into production without losing sight of them. It already runs at production scale across the Consortium.

Sovereign AI Factory

2025–Present

Building TELUS's sovereign compute infrastructure in Rimouski and the national network growing from it.

TOP500 Canada
#1
National network
4 sites
Planned scale
60K+ GPUs

The Sovereign AI Factory in Rimouski is TELUS's flagship sovereign compute facility, fully operational since September 2025 and sold out.

Rimouski is the proof point for the national expansion announced with the Government of Canada, Province of British Columbia, NVIDIA, BC Hydro, and Westbank. Kamloops, Vancouver M3, and 150 West Georgia come next. The purpose is consistent across the network: Canadian AI workloads processed, stored, and managed inside Canada.

wâsikan kisewâtisiwin

2026–Present

I led TELUS's pilot of wâsikan kisewâtisiwin, an Indigenous-built AI that detects misinformation, colonial framing, and cultural bias in writing about Indigenous Peoples.

Pilot testers
300+
Beta tester
1st

wâsikan kisewâtisiwin — "kind electricity" in Cree, named in ceremony — is Shani Gwin's company and creation: an AI trained in-house with Indigenous elders and community members that gives authenticated, Indigenous-informed guidance to anyone writing about or researching Indigenous Peoples in Canada, correcting in real time and auditing documents for bias.

The pilot kicked off in January 2026 and was announced that May, making TELUS the startup's first beta tester. I led the work: the model runs as its own deployment within Fuel iX — selectable when teams build copilots — and is hosted on the Sovereign AI Factory, keeping it on sovereign Canadian infrastructure in an OCAP®-friendly environment.

More than 300 team members joined the formal pilot, thousands more have access, and I proposed the joint purple-team effort helping the wâsikan team harden the model ahead of its public release.

Built to lift a burden, not extract knowledge

Shani Gwin, a Métis entrepreneur from Edmonton, spent years in government communications as one of the only Indigenous people on her team. Reviewing content for misinformation kept landing on her desk. She approached the Alberta Machine Intelligence Institute to test whether an AI could recognize harmful or misinformed writing about Indigenous Peoples, then built and trained the model with Indigenous elders and community members. It works inside a company's AI systems or as a writing add-on. It deliberately does not give away traditional teachings.

Fuel iX

2023–Present

Turned Unicorn.ai's shared backend into a governed platform for building and running AI across models, clouds, and products.

Custom copilots
53K+
Requests
~69M
Tokens processed
~1T

Unicorn.ai had proved demand inside TELUS. The next problem was fragmentation: every new AI product needed models, company context, tools, identity, moderation, observability, and cost controls while the underlying technology changed by the month. Nobody wanted to rebuild that stack once per application.

I led the technical work of separating those concerns into Fuel iX, and the contributors who had leaned into Unicorn became its first full-time team. Applications could change models or clouds behind a common gateway; access, policy, orchestration, and telemetry stayed in one runtime.

Retrieval at copilot scale

Copilots changed retrieval from a handful of shared knowledge bases into hundreds of thousands of personal, isolated ones. Five vector-database approaches fell short before the team moved to an object-storage-native design. It now spans more than 300,000 namespaces and 200 million documents while keeping p99 retrieval below 100 milliseconds.

Replacing LiteLLM

LiteLLM helped us move quickly, but its proxy stopped fitting the platform Fuel iX was becoming. BasicLLM replaced it with an OpenAI-compatible gateway we could shape ourselves, giving products a stable boundary while models and providers kept changing.

Anyone could build a copilot

Fuel iX Copilots let people connect private knowledge and build assistants without waiting for an AI product team. They created more than 53,000 copilots for work including HR policy, translation, brand guidance, conversation coaching, IT service, field operations, retail, and customer support. Prisms extended the same idea with conversational creation and live rendered and code views.

Governance ran with the workload

Role-based access, moderation, data sovereignty, usage and cost visibility, safety testing, and auditability lived in the runtime rather than in a review after deployment. That work was later productized as Fuel iX Fortify for automated red teaming and continuous vulnerability monitoring. A customer-support system built on the platform became the first GenAI-powered application to receive ISO 31700-1 Privacy by Design certification.

First customer: TELUS

TELUS was customer zero. It used the platform to learn what survived enterprise scale before offering it to others. Fuel iX now anchors a commercial portfolio spanning Platform, Copilots, Fortify, Agent Trainer, and Agent Assist. Across TELUS's broader AI program, more than 66,000 team members report saving an average of 3.8 hours each week, with more than $600 million in AI and advanced-analytics benefits since 2023.

Unicorn.ai

2023

Started TELUS's generative AI program with a weekend hack that grew into the shared platform behind the company's first wave of enterprise AI.

First build
65 lines
In production
Day one
Commits
2.5K+
Contributors
61

Unicorn.ai began on March 26, 2023 as 65 lines I wrote over a weekend. By the end of the day it was running in Slack and on the web. What made it interesting wasn't the chat box. It was where the AI sat: inside conversations where work was already happening, reading the same thread as everyone else, with its knowledge, costs, and actions in the open.

I developed it in the open and pushed code daily, partly to build the thing and partly to show how fast we could ship inside a company of TELUS's size.

A product before a program

Over the next two months, the hack became a Slack-native product for direct and group conversations. It learned to queue requests inside active threads, carry memories and runtime guidance, expose model choice, cost, tokens, and latency, collect feedback, and retrieve internal knowledge with citations. People from my team started lending their hands to something they could already use.

The brain became bigger than the bot

On May 24, my rewrite moved shared capabilities behind a common backend, and Unicorn.brain was initialized the same day. Sara Ghaemi made that first repository commit and led substantial early implementation while I continued shaping the product and platform direction. Slack and web became clients of shared APIs that could also serve call-centre tools and customer experiences. Nate Axcell's SimplifyHub on Backstage became the longer-term web home.

The Google Chat client

With the backend split out, a third client was an easy entry. Beginning that September, people from my org—no assigned team—built the Google Chat client end to end: authentication, direct messages and shared spaces, thread history, Google Drive context, streaming, telemetry, copilots, and feedback, all through Unicorn.brain's APIs. I didn't have to write any of it.

Achievements

  • Slack and web moved behind a shared API within weeks of the platform split
  • Google Chat joined Slack and web as the platform's third client

GitHub Enterprise

2022

Consolidated fragmented developer tooling onto GitHub and built security and automation into the everyday workflow.

Cost savings
$16.9M
Developers
4000+
Survey time saved
2h/week

Katie Peters, a staff developer on my team, championed a clear vision: focus, flow, and joy for developers. As head of engineering productivity, I led the enterprise adoption that made GitHub the single source of truth.

We rolled out push protection enterprise-wide in one day. It was non-blocking, allowing developers to acknowledge a finding and continue. We received zero complaints. Secret scanning remediated thousands of accidentally shared secrets, while GitHub Actions automated testing and validation.

Achievements

  • Dependabot monitoring dependencies enterprise-wide
  • Deployments decoupled from releases

TunnelBear

Head of Engineering

2019–2021Toronto, Canada

I ran engineering at TunnelBear through its first years inside McAfee, leading its consumer products, global network, partner platform, and customer support. Monthly active users more than doubled in six months while we kept the privacy promises and the personality that people trusted TunnelBear for.

Selected workView

VPN for Millions

TunnelBear's VPN absorbed a sudden distribution-driven surge while we improved its visibility, capacity, and resilience.

Monthly users
~8.5M
People
~40

When McAfee added VPN to its core protection products, the network did not get a gentle ramp. We had to improve the experience while millions of people were already relying on it.

Adding servers was the easy part. We needed to see each connection, distinguish provider issues from client issues, and make live network health part of incident response.

The 40-person reset

TunnelBear's commitments had multiplied with the acquisition: the consumer VPN, RememBear, a shared platform and SDK, McAfee products, and external partners. The team had not. I reset the roughly 40-person Development, DevOps, and Support organization around the outcomes the VPN needed to deliver—connect anywhere, feel fast, and never force people to turn it off—and organized the work through capability groups and autonomous squads with the scope and sponsorship to solve customer problems end to end.

Achievements

  • Aligned Connectivity, Protocols, Discoverability, and Support around three customer outcomes
  • Largest bare-metal expansion in company history launched
  • Live connection and network health across clients and partners
  • PolarBear backend moved to AWS with zero planned downtime

PolarBear Platform

Built PolarBear into the platform layer carrying TunnelBear's VPN technology into McAfee and partner products.

Partner products
16+
Networks
2

McAfee did not acquire TunnelBear only for the bear. The larger opportunity was to bring its simple approach to privacy into many more products. PolarBear turned a consumer VPN system into a reusable capability that partners could integrate.

We designed partner enablement as a repeatable path from SDK and API to shared backend and self-service, while preserving the isolation required by regulated products.

Achievements

  • Adopted across partner products and distribution relationships
  • Partner SDK moved from plan to shipped platform in 2019
  • Developer Portal made integrations increasingly self-service

Outsmarting Censorship

TunnelBear had always scrambled when a government cut the internet; we made keeping people connected a standing program.

Iran allowance
10GB
Civil-society accounts
20K

TunnelBear had long responded to shutdowns by opening free access wherever people were being silenced. We made that instinct systematic: a dedicated anti-censorship team and a four-stage framework for getting the app into a country, reaching our APIs, establishing a VPN connection, and keeping it alive.

During the 2020 crises in Venezuela, Belarus, and Iran, the team combined access programs, protocol work, civil-society partnerships, and open source into a repeatable response instead of treating every shutdown as an isolated emergency.

Achievements

  • Iran became TunnelBear's first long-term country program
  • NGO Support Network for activists, journalists, and human-rights defenders
  • Encrypted SNI became a working last resort when other techniques failed
  • Encrypted SNI changes released as open source

Trust in Public

Kept TunnelBear's promises expensive: public audits, transparency reports, and pulling Hong Kong servers when the law made tomorrow's privacy impossible to guarantee.

Audit duration
40 days
Cure53 testers
9
Usage data provided
0
Server removal
<2 weeks

A VPN sits between people and the internet, so asking for blind trust is not good enough. TunnelBear had established the industry's first public independent security audit before I arrived. I kept that commitment as the surface expanded across apps, infrastructure, AWS, browser extensions, websites, and PolarBear.

The fourth annual review covered the enlarged codebase and operating surface. The team remediated its findings and published the report; the transparency report applied the same standard to government and law-enforcement requests by stating precisely what TunnelBear could and could not provide.

Privacy as an infrastructure decision

Hong Kong's national security law changed the risk model overnight. We did not store personally identifiable information on our servers there; the danger was that state control of hardware and configuration keys could turn encryption that was safe today into exposure tomorrow. We expanded Singapore and Japan capacity first, disabled the physical Hong Kong servers within two weeks of the law, and defended the decision publicly. People in Hong Kong never lost service.

Achievements

  • Fourth annual independent audit published in full
  • Physical Hong Kong servers disabled within two weeks of the law
  • White-box review across applications, code, and infrastructure
  • No usage data provided in response to 22 requests during 2020

Loblaw Digital

Director of Engineering

2015–2019Toronto, Canada

I joined Loblaw Digital in 2015 to automate testing, then founded Engineering Productivity. As its work expanded across internal tools, delivery, platforms, mobile, and reliability, my role changed from building the first answer to building the teams and leaders who could run with it. By 2019, the organization had grown to 65 people serving the whole business.

Selected workView

Engineering Productivity

Quality, delivery, infrastructure, and reliability treated as one connected system from code to production.

Open-source libraries
33
Talks
36
Published articles
9

Quality was a downstream gate when I arrived: a small testing group, long feedback cycles, and a growing browser suite expected to catch problems after development. Automating that queue would only have made it bigger. We moved quality into development, then followed the constraints that surfaced in tooling, delivery, infrastructure, and production.

Quality moved into development

By 2016, Test Engineering was organized across grocery, fulfilment, stores, mobile, apparel, beauty, and pharmacy, backed by shared Internal Tooling. Its mission was specific: give teams maintainable frameworks they could extend, detect failures early, stop broken source before it moved downstream, and do it without turning every feature engineer into a test-infrastructure specialist.

Sunsetting QA and DevOps

By 2019, the organization had explicit Test, Tools, Release, and Platform capabilities alongside Internal Tools, Mobile Productivity, and Site Reliability Engineering. Its plan called for sunsetting both QA and DevOps as separately staffed functions. Known application patterns generated their own build, test, and deployment pipelines; feature flags and canaries separated deployment from release; infrastructure, policy, health, and observability travelled with the service.

Reduce Toil. Increase Happiness. Get Shit Done.

New ideas usually started with one person building the first version, then spread through guilds and core followers; the originator didn't own them forever. We judged internal tools the way you'd judge a gift: did it make someone else's work better? Richard Song wrote it all down in 2019.

Achievements

  • Common build, test, and deployment paths covered static sites, services, iOS, and Android
  • The engineering blog, open-source work, HashiCorp and Ruby testing meetups, and internal HackDays gave teams places to teach and extend what they learned

Bueller and Ferris

One reusable grocery suite, then a family of testing products maintained across Loblaw Digital.

Second banner
6 days
Tests per day
250K
Concurrent builds
20
Regression cycle reduction
2.5×

Six days after I started Bueller, the same grocery suite ran successfully against a second Loblaw banner by changing only its root URL. That small result proved the design: the customer journey could remain the same while the brand, language, environment, browser, and device changed underneath it.

Bueller became Ferris and expanded beyond grocery into fashion, beauty, loyalty, and pharmacy. The common machinery was extracted into a shared library while individual products developed their own suites and maintainers. The practice moved below the browser as well, adding reusable API clients and self-service tools for creating users, carts, orders, and other product state.

The bigger change was ownership. Other engineers took responsibility for individual products, extended the shared system, and replaced parts as the technology changed. I spent less time writing automation and more time building the teams around it.

Achievements

  • One grocery suite covered accounts, carts, checkout, deals, recipes, lists, payments, localization, and delivery timeslots across four retail banners
  • Desktop, tablet, and mobile behaviour selected at runtime, with SMS delivery verified through a real phone
  • Automated results flowed into reports, JIRA, team messaging, and InfluxDB rather than living in isolated test runs

Cloud Platform

Cloud infrastructure as a shared product for teams building and operating software at Loblaw Digital.

Grocery migration
6 months
Grocery performance
Grocery capacity
Grocery compute
−33%

When I founded Platform Engineering, Loblaw Digital's products were split between data-centre infrastructure and a growing collection of cloud experiments. We built a common operating platform so product teams did not each have to invent how to provision, deploy, secure, and run their services.

The first major proof was online grocery. Working with Google Cloud Professional Services and Publicis Sapient, we moved its existing SAP Hybris system to Google Cloud in six months. The initial lift-and-shift was deliberately pragmatic, while its Terraform foundations and architecture set up the later move to Kubernetes.

That production migration made cloud part of the normal path from a code change to an owned service. Platform Engineering connected infrastructure, delivery, observability, and reliability around the teams responsible for the products.

Digital Experience Studio

The path from idea to live customer experience, rebuilt without application servers or infrastructure handoffs.

Time to interactive
9.4× faster
Performance
+92%
Accessibility
+26%
Monthly savings
$38K

Loblaw Digital had more than 150 web experiences to modernize, but creating one meant navigating Windows servers, central infrastructure, and hand-operated releases. I co-founded the Digital Experience Studio to replace that delivery path, not simply push the same backlog through it faster.

Small, autonomous teams brought product, design, content, and engineering together around reusable components, starter kits, and a shared delivery platform. Preview environments and static delivery on Netlify removed the application-server queue, so a team could take an experience from idea to production on its own.

Achievements

  • 11 products delivered in four months, versus nine in the preceding seven years
  • Sites and campaigns launched in minutes through automated preview and delivery
  • Zero application servers in the new delivery model

Iris

Made internal systems usable from Slack by the people closest to the work.

Routine tasks
1–2 hours → seconds
Saved in 7 weeks
43.5 hours

Iris grew out of Jeanie, a Ruby bot I built after noticing that PC Express colleagues were already requesting operational changes in Slack, then waiting for someone with access to the right system to carry them out. Jeanie proved the idea; Iris was the Internal Tools team's ground-up rebuild in Elixir.

The team turned a collection of commands into a permission-aware platform with interactive forms, multi-step approvals, persistent sessions, observability, and a modular architecture other development teams could extend.

Through Iris, field colleagues could correct product listings and adjust store operations from their phones. Employees could request access, while engineering teams could prepare releases and coordinate incidents without manual handoffs.

Achievements

  • A mislabeled product could be removed from the online store in seconds rather than waiting hours
  • Release notes were assembled from code changes and JIRA work, then published automatically
  • Access requests carried submission, approval, fulfillment, and feedback through one workflow

From the Loblaw team

This wonderful human being was my first manager in software coming out of school, and he was instrumental in setting the standards, tone and philosophy for the rest of my career.

Richard Song
01 / 13

Redknee

Sr. Systems Engineer

2010–2015Global

I spent five years putting the billing, charging, and customer-care systems behind live mobile services into production in more than 25 countries. Repeating that work across carriers and languages taught me to turn field problems into reusable systems, and eventually it pulled me out of field engineering altogether.

Selected workView

Global Deployments

Made mediation, rating, charging, provisioning, customer care, and local operations work as one carrier system.

Each carrier implementation crossed the full path from a network event to a customer bill: mediation, rating, charging, provisioning, customer care, and the operational systems around them. The product stayed recognizable, but every market brought different infrastructure, integrations, languages, launch teams, and operating constraints.

That repetition exposed the real cost of bespoke delivery. Too much of every implementation was rebuilt or verified by hand. Fizz, and later the hosted platform, came out of that.

Achievements

  • Kuwait: First major international deployment
  • Implementations across Botswana, Zimbabwe, Fiji, Vanuatu, Tonga, Kyrgyzstan, Germany, and Ireland

Fizz

Encoded an end-to-end carrier launch in software.

Acceptance suites
43
Carrier and environment configurations
26
Carrier brands
4
Languages
3

For five years, I watched carrier implementations scale the same way: add people, fly them to the customer, and spend weeks or months working through the system by hand. Activating a subscriber was never one task. It crossed customer and dealer portals, back-office billing, APIs, command-line tools, email, and the phone itself. When the next customer or release arrived, much of that work began again.

Jeff Morgan's Cucumber & Cheese supplied the missing idea: the process could be expressed as readable software. I built Fizz around scenarios that delivery teams could read, then performed those scenarios across the entire system. Carrier, language, browser, release, and environment details were kept outside the workflow, allowing the same acceptance knowledge to be reused instead of reconstructed for every implementation.

Fizz changed more than the speed of testing. A process that had depended on a large group being on site could now be repeated, run in parallel, inspected, handed over, and extended as the product changed. It was the first time I turned what I'd learned in the field into a system that could travel instead of the team.

Achievements

  • Covered activation, plans, payments, billing, add-ons, usage, SIM and phone-number changes, voicemail, and account care
  • Created and managed the accounts, vouchers, and SIMs required by each run, including records for later production cleanup
  • Generated requirements and handover documentation from the same scenarios it executed
  • Produced reports with screenshots, diagnostics, timings, and environment data, then returned results through JIRA and email

Americas Cloud Platform

A managed billing and customer-care platform that made new carrier launches repeatable.

Hosted onboarding
2 weeks
Legacy replacement
13 weeks
MVNOs supported
5

After years implementing the same stack customer by customer, I implemented and launched Redknee's Americas cloud platform as a managed service. It brought activation, provisioning, real-time rating and charging, customer care, dealer tools, and self-service into one operating model, so a new carrier could be onboarded without rebuilding the delivery machinery each time.

Fizz was part of the connective tissue that kept acceptance work portable across brands and environments.

Education

Degree

University of Toronto

Honours Bachelor · Anthropology & Computer Science

Semiotics Researcher, Computer Science Student Union Social Director

2005–2010

Program

2010

Arctic Field School

University of Manitoba · Ecology, History, and Land Use

In Pangnirtung, I studied Inuktitut, Inuit history, ecology, and local land use.

Certificate

2025

TELUS Senior Leadership Forum

Massachusetts Institute of Technology · Sloan School of Management