Kubernetes
The runtime platform where our container becomes a resilient service.
How This Fits Our End-to-End Pipeline
Kubernetes is documented here as part of one connected build, not as an isolated tutorial. The goal is to explain how code moved from a developer laptop into a monitored Kubernetes service.
Developer laptop
| git add / commit / push
v
GitHub repository
| pull request + branch rules
v
GitHub Actions workflow
| test -> docker build -> tag -> push
v
Docker Hub image registry
| immutable image tag
v
GitOps manifests repository/path
| ArgoCD watches desired state
v
Kubernetes cluster
| app pods + services + ingress
v
Prometheus scrapes metrics -> Grafana dashboards
Theory
Kubernetes runs containers across a cluster. Instead of starting containers manually, we declare the desired state using YAML: deployments define pods, services provide stable networking, config maps and secrets inject settings, ingress exposes HTTP routes, and namespaces separate environments.
In our pipeline, Kubernetes is the runtime layer. GitHub Actions does not SSH into servers. It builds an image and updates manifests. ArgoCD applies those manifests, and Kubernetes reconciles the cluster until reality matches the desired state.
Cluster Architecture
User -> Ingress Controller -> Service -> Deployment -> ReplicaSet -> Pods -> Containers
| | |
| | +-- ConfigMap / Secret mounted or env injected
| +-- Rolling update strategy controls pod replacement
+-- DNS and service discovery hide changing pod IPs
Deployments
A Deployment manages replicas and rolling updates for stateless workloads. It owns ReplicaSets, which own Pods. We change the image tag in the Deployment manifest, and Kubernetes safely replaces old pods with new ones.
apiVersion: apps/v1
kind: Deployment
metadata:
name: devops-demo
namespace: staging
spec:
replicas: 2
strategy:
type: RollingUpdate
rollingUpdate:
maxUnavailable: 0
maxSurge: 1
selector:
matchLabels:
app: devops-demo
template:
metadata:
labels:
app: devops-demo
spec:
containers:
- name: app
image: dockerhub-user/devops-demo:sha-9f3a21c
ports:
- containerPort: 3000
readinessProbe:
httpGet:
path: /health
port: 3000
livenessProbe:
httpGet:
path: /health
port: 3000
Services
Pods are temporary and receive changing IP addresses. A Service gives a stable DNS name and virtual IP for reaching matching pods.
apiVersion: v1
kind: Service
metadata:
name: devops-demo
namespace: staging
spec:
type: ClusterIP
selector:
app: devops-demo
ports:
- port: 80
targetPort: 3000
ConfigMaps and Secrets
ConfigMaps hold non-sensitive settings. Secrets hold sensitive values but should still be treated carefully because basic Kubernetes Secrets are base64 encoded, not automatically encrypted for every use case.
apiVersion: v1
kind: ConfigMap
metadata:
name: devops-demo-config
namespace: staging
data:
APP_ENV: staging
LOG_LEVEL: info
---
apiVersion: v1
kind: Secret
metadata:
name: devops-demo-secret
namespace: staging
type: Opaque
stringData:
API_TOKEN: replace-with-secret-manager-output
Namespaces
- Use
stagingandproductionnamespaces to separate environments inside one cluster. - Apply resource quotas and network policies per namespace as the setup matures.
- Avoid deploying everything into
default; it becomes messy quickly.
kubectl create namespace staging
kubectl config set-context --current --namespace=staging
Rolling Updates
Rolling updates replace pods gradually. Readiness probes are the safety signal: a new pod should not receive traffic until it is ready.
kubectl set image deployment/devops-demo app=dockerhub-user/devops-demo:sha-new
kubectl rollout status deployment/devops-demo
kubectl rollout undo deployment/devops-demo
Ingress Concepts
Ingress routes external HTTP traffic to internal services. It normally requires an ingress controller such as NGINX Ingress Controller, Traefik, or a cloud provider controller.
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
name: devops-demo
namespace: staging
spec:
rules:
- host: devops-demo.local
http:
paths:
- path: /
pathType: Prefix
backend:
service:
name: devops-demo
port:
number: 80
Commands We Used
Inspect cluster
kubectl cluster-info
kubectl get nodes
Confirms kubectl can talk to the cluster.
Apply manifests
kubectl apply -f k8s/
Creates or updates resources from YAML files.
List resources
kubectl get all -n staging
Shows pods, services, deployments, and replicasets in the namespace.
Describe a pod
kubectl describe pod <pod-name> -n staging
Shows scheduling events, image pull errors, probe failures, and container state.
Read logs
kubectl logs deployment/devops-demo -n staging
Reads app logs through the deployment selector.
Debug networking
kubectl port-forward service/devops-demo 8080:80 -n staging
Forwards local port 8080 to the service inside the cluster.
Troubleshooting Flow
- Start with
kubectl get pods -n staging. - If the pod is pending, describe it and inspect scheduling events.
- If it is waiting, check image pull errors and secret configuration.
- If it is running but not ready, inspect readiness probes and app logs.
- If service traffic fails, compare service selectors with pod labels.
Quiz Questions
Q1. Why do we need a Service if a Deployment already creates pods?
Answer: Because pods have unstable IPs; a Service gives a stable endpoint and load balances to matching pods.
Q2. What happens when a readiness probe fails?
Answer: The pod stays out of service endpoints, so it should not receive traffic.
Q3. Why should image tags be immutable?
Answer: So the running workload can be traced back to an exact build and commit.
Q4. What is the difference between ConfigMap and Secret?
Answer: ConfigMap is for non-sensitive configuration; Secret is for sensitive data, though it still needs careful storage and access control.
Q5. What does ArgoCD add on top of kubectl apply?
Answer: Continuous reconciliation, drift detection, sync status, history, and a GitOps operating model.
Interview Notes
- Deployment manages stateless replicated pods; StatefulSet manages stable identity and storage for stateful apps.
- ClusterIP exposes internally, NodePort exposes on node ports, LoadBalancer asks infrastructure for an external load balancer.
- Readiness is for traffic eligibility; liveness is for restarting unhealthy containers.
- Ingress is routing; an ingress controller is the actual component that implements it.
- Declarative YAML plus reconciliation is the core Kubernetes mental model.
Memorize vs Look-up
Memorize
kubectl get,describe,logs,apply,rollout status,rollout undo.- Deployment, Service, ConfigMap, Secret, Namespace, Ingress roles.
- Readiness versus liveness probes.
Look Up
- Every API version and manifest field.
- Cloud-provider-specific LoadBalancer annotations.
- Advanced network policy and storage class details.
Project Implementation Journal
Baseline
We first made sure the application or configuration worked before introducing Kubernetes. A broken baseline makes every later automation failure harder to understand.
Local proof
We validated the Kubernetes workflow locally where possible, because local feedback is faster than waiting for CI or a cluster reconciliation loop.
Repository proof
We committed the Kubernetes change as a reviewable unit so the reason for the pipeline change was visible in Git history.
Automation handoff
We connected Kubernetes to the next tool in the chain instead of treating it as a standalone exercise.
Failure check
We intentionally inspected the common failure signals for Kubernetes: logs, status output, permissions, names, tags, and configuration paths.
Rollback thinking
We asked how to return to the previous working state if the Kubernetes change caused a bad deployment.
Interview compression
We reduced Kubernetes into a few sentences that explain purpose, implementation, and failure modes clearly.
Operations note
We documented what someone should check the day after the deployment, not only what to run during setup.
Decision Records
- Keep Kubernetes configuration in Git where it can be reviewed.
- Prefer explicit names over clever names: repository, image, namespace, workflow, and application names should be searchable.
- Use immutable versions for anything that can be deployed or rolled back.
- Put secrets in the platform secret store, not in code, documentation screenshots, or shell history.
- Automate only after the manual path is understood.
- Make the happy path visible, then document the first five things to check when it fails.
- Separate staging and production concerns before the project becomes too large.
- Choose boring defaults unless there is a real operational reason to customize.
- Write commands so they can be pasted into a terminal after replacing obvious placeholders.
- Treat dashboards, manifests, and workflow files as production code once people rely on them.
Failure Modes We Learned To Recognize
Wrong name
The most ordinary failures came from mismatched names: image repository, namespace, service selector, branch, workflow file, or dashboard variable.
Wrong permission
Automation failed when a token could read but not write, push but not pull, or access staging but not production.
Wrong version
A deployment looked successful while the cluster still ran an old image tag or an image tag that had been overwritten.
Wrong assumption
A command that worked locally failed in CI because the runner had a different shell, path, network, or credential context.
Missing feedback
Without logs, status commands, metrics, or dashboards, the system gave no quick answer about what changed.
Manual drift
Manual cluster edits solved a momentary problem but made the GitOps source of truth inaccurate.
Interview Drill Questions
- What problem does Kubernetes solve in this pipeline?
- What artifact or state does Kubernetes produce?
- Which tool consumes the output of Kubernetes next?
- What is the most likely beginner mistake with Kubernetes?
- How would you prove Kubernetes worked without guessing?
- How would you roll back a bad change involving Kubernetes?
- What should be memorized versus looked up for Kubernetes?
- Which security boundary matters most for Kubernetes?
- How would you explain Kubernetes to someone who only knows basic Linux?
- What metric, log, status, or command would you check first during an incident?
Glossary For This Stage
Artifact
A build output or configuration object that can be handed to another stage.
Desired state
The state declared in Git or YAML that controllers try to make real.
Reconciliation
The loop where a tool compares desired state with actual state and fixes differences.
Immutable version
A version reference that should never change meaning after publication.
Rollback
A controlled return to a previously known working state.
Drift
A difference between what Git says should exist and what is actually running.
Health
A status signal that says whether the service is ready and operating correctly.
Traceability
The ability to connect a running system back to a commit, workflow run, image, and manifest change.
Practical Runbook
Confirm source
Identify the exact repository, branch, commit, file path, or dashboard connected to Kubernetes.
Confirm identity
Check the account, token, kube context, registry namespace, or runner label before assuming the tool is broken.
Confirm version
Write down the version or tag you expected and compare it with the version the platform reports.
Confirm status
Use the native status command or UI first; it usually tells you whether the failure is configuration, permission, or runtime.
Confirm logs
Logs explain what happened after the tool accepted the configuration but the process still failed.
Confirm network
Many CI/CD failures are actually DNS, registry, cluster, firewall, or service discovery failures.
Confirm ownership
Know whether the application team, platform team, security team, or repository owner controls the failing setting.
Confirm rollback
Before changing more things, decide whether the quickest safe move is to revert, resync, rebuild, or redeploy.
Confirm documentation
After fixing the issue, add the command, symptom, and fix to the project notes so the next run is faster.
Confirm automation
If the Kubernetes step is repeated manually more than twice, turn it into a workflow, manifest, script, or checklist.
Verification Checklist
- Can I point to the exact Git commit connected to this Kubernetes change?
- Can I explain what changed in one sentence?
- Can I prove the change worked with a command or status screen?
- Can I identify the next tool that consumes this output?
- Can I identify the secret, token, or permission that would break this step?
- Can I roll back without manually editing production state?
- Can I tell whether the failure is build-time, deploy-time, or runtime?
- Can I show the relevant logs or metrics?
- Can I repeat the setup on a fresh machine or cluster?
- Can I teach this stage without reading every command from the page?