ArgoCD

The GitOps controller that deploys what Git declares.

How This Fits Our End-to-End Pipeline

ArgoCD is documented here as part of one connected build, not as an isolated tutorial. The goal is to explain how code moved from a developer laptop into a monitored Kubernetes service.

Developer laptop
  | git add / commit / push
  v
GitHub repository
  | pull request + branch rules
  v
GitHub Actions workflow
  | test -> docker build -> tag -> push
  v
Docker Hub image registry
  | immutable image tag
  v
GitOps manifests repository/path
  | ArgoCD watches desired state
  v
Kubernetes cluster
  | app pods + services + ingress
  v
Prometheus scrapes metrics -> Grafana dashboards

Theory

ArgoCD is a GitOps controller for Kubernetes. It watches Git repositories, compares the desired manifests with the live cluster, reports drift, and syncs changes into Kubernetes.

Our pipeline uses ArgoCD as the deployment engine. GitHub Actions builds and pushes the image, then updates the manifest tag. ArgoCD notices the Git change and applies it to the cluster.

GitOps Diagram

Git repository path: k8s/staging
        |
        | ArgoCD watches branch + path
        v
ArgoCD Application
        |
        | compare desired vs live
        v
Kubernetes API
        |
        | create/update resources
        v
Pods, Services, ConfigMaps, Secrets, Ingress

Application Manifest

apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
  name: devops-demo-staging
  namespace: argocd
spec:
  project: default
  source:
    repoURL: https://github.com/<user>/<repo>.git
    targetRevision: main
    path: k8s/staging
  destination:
    server: https://kubernetes.default.svc
    namespace: staging
  syncPolicy:
    automated:
      prune: true
      selfHeal: true
    syncOptions:
      - CreateNamespace=true

Step-by-step Implementation

  1. Install ArgoCD into an argocd namespace.
  2. Create a Kubernetes manifests folder such as k8s/staging.
  3. Create an ArgoCD Application that points to the repository path.
  4. Let GitHub Actions update image tags in that path after a successful build.
  5. Use the ArgoCD UI or CLI to watch sync status, health, and drift.

Commands

Install ArgoCD

kubectl create namespace argocd
kubectl apply -n argocd -f https://raw.githubusercontent.com/argoproj/argo-cd/stable/manifests/install.yaml

Installs ArgoCD controllers and API server into the cluster.

Port forward UI

kubectl port-forward svc/argocd-server -n argocd 8080:443

Exposes the ArgoCD UI locally for learning setups.

Apply application

kubectl apply -f argocd/application.yaml

Registers the app that ArgoCD should manage.

Check status

argocd app get devops-demo-staging
argocd app sync devops-demo-staging

Shows and manually syncs application state when using the CLI.

Drift Example

If someone manually edits a Deployment with kubectl edit, ArgoCD can show the app as OutOfSync because live state no longer matches Git. With self-heal enabled, ArgoCD restores the Git version.

Promotion Model

  • Staging auto-syncs from main.
  • Production syncs from a release branch or tagged manifest update.
  • Promotion is a Git change, not an SSH session.

Interview Notes

  • GitOps means Git is the desired-state source, while controllers reconcile the live system to match it.
  • ArgoCD detects drift when the cluster differs from Git.
  • Self-heal restores live state to Git; prune removes resources no longer defined in Git.
  • ArgoCD sync is pull-based from the cluster perspective, which reduces the need to give CI direct cluster credentials.
  • Application, Project, Repository, Cluster, Sync, Health, and Drift are key ArgoCD terms.

FAQs

Does ArgoCD build Docker images?

No. It deploys Kubernetes manifests. Image building belongs in CI, such as GitHub Actions.

Should auto-sync be enabled?

For staging, usually yes. For production, teams may require manual approval or progressive delivery.

What is OutOfSync?

The desired state in Git and the live state in Kubernetes do not match.

Memorize vs Look-up

Memorize

  • GitOps desired state model.
  • OutOfSync, Synced, Healthy, Degraded meanings.
  • Prune removes old resources; self-heal repairs drift.

Look Up

  • All ArgoCD RBAC policy syntax.
  • SSO connector configuration.
  • Advanced sync waves and hooks.

Project Implementation Journal

Baseline

We first made sure the application or configuration worked before introducing ArgoCD. A broken baseline makes every later automation failure harder to understand.

Local proof

We validated the ArgoCD workflow locally where possible, because local feedback is faster than waiting for CI or a cluster reconciliation loop.

Repository proof

We committed the ArgoCD change as a reviewable unit so the reason for the pipeline change was visible in Git history.

Automation handoff

We connected ArgoCD to the next tool in the chain instead of treating it as a standalone exercise.

Failure check

We intentionally inspected the common failure signals for ArgoCD: logs, status output, permissions, names, tags, and configuration paths.

Rollback thinking

We asked how to return to the previous working state if the ArgoCD change caused a bad deployment.

Interview compression

We reduced ArgoCD into a few sentences that explain purpose, implementation, and failure modes clearly.

Operations note

We documented what someone should check the day after the deployment, not only what to run during setup.

Decision Records

  • Keep ArgoCD configuration in Git where it can be reviewed.
  • Prefer explicit names over clever names: repository, image, namespace, workflow, and application names should be searchable.
  • Use immutable versions for anything that can be deployed or rolled back.
  • Put secrets in the platform secret store, not in code, documentation screenshots, or shell history.
  • Automate only after the manual path is understood.
  • Make the happy path visible, then document the first five things to check when it fails.
  • Separate staging and production concerns before the project becomes too large.
  • Choose boring defaults unless there is a real operational reason to customize.
  • Write commands so they can be pasted into a terminal after replacing obvious placeholders.
  • Treat dashboards, manifests, and workflow files as production code once people rely on them.

Failure Modes We Learned To Recognize

Wrong name

The most ordinary failures came from mismatched names: image repository, namespace, service selector, branch, workflow file, or dashboard variable.

Wrong permission

Automation failed when a token could read but not write, push but not pull, or access staging but not production.

Wrong version

A deployment looked successful while the cluster still ran an old image tag or an image tag that had been overwritten.

Wrong assumption

A command that worked locally failed in CI because the runner had a different shell, path, network, or credential context.

Missing feedback

Without logs, status commands, metrics, or dashboards, the system gave no quick answer about what changed.

Manual drift

Manual cluster edits solved a momentary problem but made the GitOps source of truth inaccurate.

Interview Drill Questions

  • What problem does ArgoCD solve in this pipeline?
  • What artifact or state does ArgoCD produce?
  • Which tool consumes the output of ArgoCD next?
  • What is the most likely beginner mistake with ArgoCD?
  • How would you prove ArgoCD worked without guessing?
  • How would you roll back a bad change involving ArgoCD?
  • What should be memorized versus looked up for ArgoCD?
  • Which security boundary matters most for ArgoCD?
  • How would you explain ArgoCD to someone who only knows basic Linux?
  • What metric, log, status, or command would you check first during an incident?

Glossary For This Stage

Artifact

A build output or configuration object that can be handed to another stage.

Desired state

The state declared in Git or YAML that controllers try to make real.

Reconciliation

The loop where a tool compares desired state with actual state and fixes differences.

Immutable version

A version reference that should never change meaning after publication.

Rollback

A controlled return to a previously known working state.

Drift

A difference between what Git says should exist and what is actually running.

Health

A status signal that says whether the service is ready and operating correctly.

Traceability

The ability to connect a running system back to a commit, workflow run, image, and manifest change.

Practical Runbook

Confirm source

Identify the exact repository, branch, commit, file path, or dashboard connected to ArgoCD.

Confirm identity

Check the account, token, kube context, registry namespace, or runner label before assuming the tool is broken.

Confirm version

Write down the version or tag you expected and compare it with the version the platform reports.

Confirm status

Use the native status command or UI first; it usually tells you whether the failure is configuration, permission, or runtime.

Confirm logs

Logs explain what happened after the tool accepted the configuration but the process still failed.

Confirm network

Many CI/CD failures are actually DNS, registry, cluster, firewall, or service discovery failures.

Confirm ownership

Know whether the application team, platform team, security team, or repository owner controls the failing setting.

Confirm rollback

Before changing more things, decide whether the quickest safe move is to revert, resync, rebuild, or redeploy.

Confirm documentation

After fixing the issue, add the command, symptom, and fix to the project notes so the next run is faster.

Confirm automation

If the ArgoCD step is repeated manually more than twice, turn it into a workflow, manifest, script, or checklist.

Verification Checklist

  • Can I point to the exact Git commit connected to this ArgoCD change?
  • Can I explain what changed in one sentence?
  • Can I prove the change worked with a command or status screen?
  • Can I identify the next tool that consumes this output?
  • Can I identify the secret, token, or permission that would break this step?
  • Can I roll back without manually editing production state?
  • Can I tell whether the failure is build-time, deploy-time, or runtime?
  • Can I show the relevant logs or metrics?
  • Can I repeat the setup on a fresh machine or cluster?
  • Can I teach this stage without reading every command from the page?