diff --git a/docs/Makefile b/docs/Makefile new file mode 100644 index 00000000..dabd9922 --- /dev/null +++ b/docs/Makefile @@ -0,0 +1,3 @@ +workshop.md: workshop.yml *.md + ./markmaker.py < workshop.yml > workshop.md + # open http://localhost:8888/workshop.html diff --git a/docs/creatingswarm.md b/docs/creatingswarm.md new file mode 100644 index 00000000..f97e2265 --- /dev/null +++ b/docs/creatingswarm.md @@ -0,0 +1,346 @@ +# Creating our first Swarm + +- The cluster is initialized with `docker swarm init` + +- This should be executed on a first, seed node + +- .warning[DO NOT execute `docker swarm init` on multiple nodes!] + + You would have multiple disjoint clusters. + +.exercise[ + +- Create our cluster from node1: + ```bash + docker swarm init + ``` + +] + +-- + +class: advertise-addr + +If Docker tells you that it `could not choose an IP address to advertise`, see next slide! + +--- + +class: advertise-addr + +## IP address to advertise + +- When running in Swarm mode, each node *advertises* its address to the others +
+ (i.e. it tells them *"you can contact me on 10.1.2.3:2377"*) + +- If the node has only one IP address (other than 127.0.0.1), it is used automatically + +- If the node has multiple IP addresses, you **must** specify which one to use +
+ (Docker refuses to pick one randomly) + +- You can specify an IP address or an interface name +
(in the latter case, Docker will read the IP address of the interface and use it) + +- You can also specify a port number +
(otherwise, the default port 2377 will be used) + +--- + +class: advertise-addr + +## Which IP address should be advertised? + +- If your nodes have only one IP address, it's safe to let autodetection do the job + + .small[(Except if your instances have different private and public addresses, e.g. + on EC2, and you are building a Swarm involving nodes inside and outside the + private network: then you should advertise the public address.)] + +- If your nodes have multiple IP addresses, pick an address which is reachable + *by every other node* of the Swarm + +- If you are using [play-with-docker](http://play-with-docker.com/), use the IP + address shown next to the node name + + .small[(This is the address of your node on your private internal overlay network. + The other address that you might see is the address of your node on the + `docker_gwbridge` network, which is used for outbound traffic.)] + +Examples: + +```bash +docker swarm init --advertise-addr 10.0.9.2 +docker swarm init --advertise-addr eth0:7777 +``` + +--- + +class: extra-details + +## Using a separate interface for the data path + +- You can use different interfaces (or IP addresses) for control and data + +- You set the _control plane path_ with `--advertise-addr` + + (This will be used for SwarmKit manager/worker communication, leader election, etc.) + +- You set the _data plane path_ with `--data-path-addr` + + (This will be used for traffic between containers) + +- Both flags can accept either an IP address, or an interface name + + (When specifying an interface name, Docker will use its first IP address) + +--- + +## Token generation + +- In the output of `docker swarm init`, we have a message + confirming that our node is now the (single) manager: + + ``` + Swarm initialized: current node (8jud...) is now a manager. + ``` + +- Docker generated two security tokens (like passphrases or passwords) for our cluster + +- The CLI shows us the command to use on other nodes to add them to the cluster using the "worker" + security token: + + ``` + To add a worker to this swarm, run the following command: + docker swarm join \ + --token SWMTKN-1-59fl4ak4nqjmao1ofttrc4eprhrola2l87... \ + 172.31.4.182:2377 + ``` + +--- + +class: extra-details + +## Checking that Swarm mode is enabled + +.exercise[ + +- Run the traditional `docker info` command: + ```bash + docker info + ``` + +] + +The output should include: + +``` +Swarm: active + NodeID: 8jud7o8dax3zxbags3f8yox4b + Is Manager: true + ClusterID: 2vcw2oa9rjps3a24m91xhvv0c + ... +``` + +--- + +## Running our first Swarm mode command + +- Let's retry the exact same command as earlier + +.exercise[ + +- List the nodes (well, the only node) of our cluster: + ```bash + docker node ls + ``` + +] + +The output should look like the following: +``` +ID HOSTNAME STATUS AVAILABILITY MANAGER STATUS +8jud...ox4b * node1 Ready Active Leader +``` + +--- + +## Adding nodes to the Swarm + +- A cluster with one node is not a lot of fun + +- Let's add `node2`! + +- We need the token that was shown earlier + +-- + +- You wrote it down, right? + +-- + +- Don't panic, we can easily see it again 😏 + +--- + +## Adding nodes to the Swarm + +.exercise[ + +- Show the token again: + ```bash + docker swarm join-token worker + ``` + +- Switch to `node2` + +- Copy-paste the `docker swarm join ...` command +
(that was displayed just before) + +] + +--- + +class: extra-details + +## Check that the node was added correctly + +- Stay on `node2` for now! + +.exercise[ + +- We can still use `docker info` to verify that the node is part of the Swarm: + ```bash + docker info | grep ^Swarm + ``` + +] + +- However, Swarm commands will not work; try, for instance: + ``` + docker node ls + ``` + +- This is because the node that we added is currently a *worker* + +- Only *managers* can accept Swarm-specific commands + +--- + +## View our two-node cluster + +- Let's go back to `node1` and see what our cluster looks like + +.exercise[ + +- Switch back to `node1` + +- View the cluster from `node1`, which is a manager: + ```bash + docker node ls + ``` + +] + +The output should be similar to the following: +``` +ID HOSTNAME STATUS AVAILABILITY MANAGER STATUS +8jud...ox4b * node1 Ready Active Leader +ehb0...4fvx node2 Ready Active +``` + + +--- + +class: under-the-hood + +## Under the hood: docker swarm init + +When we do `docker swarm init`: + +- a keypair is created for the root CA of our Swarm + +- a keypair is created for the first node + +- a certificate is issued for this node + +- the join tokens are created + +--- + +class: under-the-hood + +## Under the hood: join tokens + +There is one token to *join as a worker*, and another to *join as a manager*. + +The join tokens have two parts: + +- a secret key (preventing unauthorized nodes from joining) + +- a fingerprint of the root CA certificate (preventing MITM attacks) + +If a token is compromised, it can be rotated instantly with: +``` +docker swarm join-token --rotate +``` + +--- + +class: under-the-hood + +## Under the hood: docker swarm join + +When a node joins the Swarm: + +- it is issued its own keypair, signed by the root CA + +- if the node is a manager: + + - it joins the Raft consensus + - it connects to the current leader + - it accepts connections from worker nodes + +- if the node is a worker: + + - it connects to one of the managers (leader or follower) + +--- + +class: under-the-hood + +## Under the hood: cluster communication + +- The *control plane* is encrypted with AES-GCM; keys are rotated every 12 hours + +- Authentication is done with mutual TLS; certificates are rotated every 90 days + + (`docker swarm update` allows to change this delay or to use an external CA) + +- The *data plane* (communication between containers) is not encrypted by default + + (but this can be activated on a by-network basis, using IPSEC, + leveraging hardware crypto if available) + +--- + +class: under-the-hood + +## Under the hood: I want to know more! + +Revisit SwarmKit concepts: + +- Docker 1.12 Swarm Mode Deep Dive Part 1: Topology + ([video](https://www.youtube.com/watch?v=dooPhkXT9yI)) + +- Docker 1.12 Swarm Mode Deep Dive Part 2: Orchestration + ([video](https://www.youtube.com/watch?v=_F6PSP-qhdA)) + +Some presentations from the Docker Distributed Systems Summit in Berlin: + +- Heart of the SwarmKit: Topology Management + ([slides](https://speakerdeck.com/aluzzardi/heart-of-the-swarmkit-topology-management)) + +- Heart of the SwarmKit: Store, Topology & Object Model + ([slides](http://www.slideshare.net/Docker/heart-of-the-swarmkit-store-topology-object-model)) + ([video](https://www.youtube.com/watch?v=EmePhjGnCXY)) diff --git a/docs/dockercon.yml b/docs/dockercon.yml new file mode 100644 index 00000000..4f63cde9 --- /dev/null +++ b/docs/dockercon.yml @@ -0,0 +1,130 @@ +chapters: +- | + class: title + + .small[ + + Swarm: from Zero to Hero + + .small[.small[ + + **Be kind to the WiFi!** + + *Use the 5G network* +
+ *Don't use your hotspot* +
+ *Don't stream videos from YouTube, Netflix, etc. +
(if you're bored, watch local content instead)* + + Thank you! + + ] + ] + ] + + --- + + ## Intros + + + + - Hello! I am + Jérôme ([@jpetazzo](https://twitter.com/jpetazzo)) + + -- + + - This is our collective Docker knowledge: + + ![Bell Curve](bell-curve.jpg) + + --- + + ## Agenda + + .small[ + - 09:00-09:10 Hello! + - 09:10-10:30 Part 1 + - 10:30-11:00 coffee break + - 11:00-12:30 Part 2 + - 12:30-13:30 lunch break + - 13:30-15:00 Part 3 + - 15:00-15:30 coffee break + - 15:30-17:00 Part 4 + - 17:00-18:00 Afterhours and Q&A + ] + + + + - All the content is publicly available (slides, code samples, scripts) + + Upstream URL: https://github.com/jpetazzo/orchestration-workshop + + - Feel free to interrupt for questions at any time + + - Live feedback, questions, help on [Gitter](chat) + + http://container.training/chat + +- intro.md +- | + @@TOC@@ +- - prereqs.md + - versions.md + - | + class: title + + All right! +
+ We're all set. +
+ Let's do this. + - sampleapp.md + - swarmkit.md + - creatingswarm.md + - morenodes.md +- - firstservice.md + - ourapponswarm.md + - updatingservices.md + - healthchecks.md +- - operatingswarm.md + - netshoot.md + - ipsec.md + - swarmtools.md +- - security.md + - secrets.md + - leastprivilege.md + - apiscope.md + - logging.md + - metrics.md + - stateful.md + - extratips.md + - end.md +- | + class: title + + That's all folks!
Questions? + + .small[.small[ + + Jérôme ([@jpetazzo](https://twitter.com/jpetazzo)) — [@docker](https://twitter.com/docker) + + ]] + + diff --git a/docs/end.md b/docs/end.md index 0c837e09..c1a7b9be 100644 --- a/docs/end.md +++ b/docs/end.md @@ -36,29 +36,3 @@ Reminder: there is a tag for each iteration of the content in the Github repository. It makes it easy to come back later and check what has changed since you did it! - ---- - -class: title, self-paced - -Thank you! - ---- - -class: title, in-person - -That's all folks!
Questions? - -.small[.small[ - -Jérôme ([@jpetazzo](https://twitter.com/jpetazzo)) — [@docker](https://twitter.com/docker) - -AJ ([@s0ulshake](https://twitter.com/s0ulshake)) — *For hire!* -
-`curl cv.soulshake.net` - -]] - - diff --git a/docs/firstservice.md b/docs/firstservice.md index 47945367..eeb6157b 100644 --- a/docs/firstservice.md +++ b/docs/firstservice.md @@ -24,14 +24,14 @@ (New in Docker Engine 17.05) -If you are running Docker 17.05 or later, you will see the following message: +If you are running Docker 17.05 to 17.09, you will see the following message: ``` Since --detach=false was not specified, tasks will be created in the background. In a future release, --detach=false will become the default. ``` -Let's ignore it for now; but we'll come back to it in just a few minutes! +You can ignore that for now; but we'll come back to it in just a few minutes! --- @@ -152,9 +152,9 @@ class: extra-details ] -Note: `--detach=false` will eventually become the default. +Note: with Docker Engine 17.10 and later, `--detach=false` is the default. -With older versions, you can use e.g.: `watch docker service ps ` +With versions older than 17.05, you can use e.g.: `watch docker service ps ` --- diff --git a/docs/intro.md b/docs/intro.md index f0f73f7c..b2f206ba 100644 --- a/docs/intro.md +++ b/docs/intro.md @@ -1,104 +1,3 @@ -class: title, self-paced - -Docker
Orchestration
Workshop - ---- - -class: title, in-person - -.small[ - -Deploy and scale containers with Docker native, open source orchestration - -.small[.small[ - -**Be kind to the WiFi!** - -*Use the 5G network* -
-*Don't use your hotspot* -
-*Don't stream videos from YouTube, Netflix, etc. -
(if you're bored, watch local content instead)* - -Thank you! - -] -] -] - ---- - -class: in-person - -## Intros - -- Hello! We are - AJ ([@s0ulshake](https://twitter.com/s0ulshake)) - & - Jérôme ([@jpetazzo](https://twitter.com/jpetazzo)) - --- - -class: in-person - -- This is our collective Docker knowledge: - - ![Bell Curve](bell-curve.jpg) - - - ---- - -class: in-person - -## Agenda - - - - - -- The tutorial will run from 9:00am to 12:20pm - -- This will be fast-paced, but DON'T PANIC! - -- All the content is publicly available (slides, code samples, scripts) - - Upstream URL: https://github.com/jpetazzo/orchestration-workshop - -- There will be a coffee break at 10:30am -
- (please remind me if I forget about it!) - -- Feel free to interrupt for questions at any time - -- Live feedback, questions, help on [Gitter](chat) - - http://container.training/chat - ---- - ## A brief introduction - This was initially written to support in-person, diff --git a/docs/machine.md b/docs/machine.md new file mode 100644 index 00000000..967629dd --- /dev/null +++ b/docs/machine.md @@ -0,0 +1,225 @@ +## Adding nodes using the Docker API + +- We don't have to SSH into the other nodes, we can use the Docker API + +- If you are using Play-With-Docker: + + - the nodes expose the Docker API over port 2375/tcp, without authentication + + - we will connect by setting the `DOCKER_HOST` environment variable + +- Otherwise: + + - the nodes expose the Docker API over port 2376/tcp, with TLS mutual authentication + + - we will use Docker Machine to set the correct environment variables +
(the nodes have been suitably pre-configured to be controlled through `node1`) + +--- + +# Docker Machine + +- Docker Machine has two primary uses: + + - provisioning cloud instances running the Docker Engine + + - managing local Docker VMs within e.g. VirtualBox + +- Docker Machine is purely optional + +- It makes it easy to create, upgrade, manage... Docker hosts: + + - on your favorite cloud provider + + - locally (e.g. to test clustering, or different versions) + + - across different cloud providers + +--- + +class: self-paced + +## If you're using Play-With-Docker ... + +- You won't need to use Docker Machine + +- Instead, to "talk" to another node, we'll just set `DOCKER_HOST` + +- You can skip the exercises telling you to do things with Docker Machine! + +--- + +## Docker Machine basic usage + +- We will learn two commands: + + - `docker-machine ls` (list existing hosts) + + - `docker-machine env` (switch to a specific host) + +.exercise[ + +- List configured hosts: + ```bash + docker-machine ls + ``` + +] + +You should see your 5 nodes. + +--- + +class: in-person + +## How did we make our 5 nodes show up there? + +*For the curious...* + +- This was done by our VM provisioning scripts + +- After setting up everything else, `node1` adds the 5 nodes + to the local Docker Machine configuration + (located in `$HOME/.docker/machine`) + +- Nodes are added using [Docker Machine generic driver](https://docs.docker.com/machine/drivers/generic/) + + (It skips machine provisioning and jumps straight to the configuration phase) + +- Docker Machine creates TLS certificates and deploys them to the nodes through SSH + +--- + +## Using Docker Machine to communicate with a node + +- To select a node, use `eval $(docker-machine env nodeX)` + +- This sets a number of environment variables + +- To unset these variables, use `eval $(docker-machine env -u)` + +.exercise[ + +- View the variables used by Docker Machine: + ```bash + docker-machine env node3 + ``` + +] + +(This shows which variables *would* be set by Docker Machine; but it doesn't change them.) + +--- + +## Getting the token + +- First, let's store the join token in a variable + +- This must be done from a manager + +.exercise[ + +- Make sure we talk to the local node, or `node1`: + ```bash + eval $(docker-machine env -u) + ``` + +- Get the join token: + ```bash + TOKEN=$(docker swarm join-token -q worker) + ``` + +] + +--- + +## Change the node targeted by the Docker CLI + +- We need to set the right environment variables to communicate with `node3` + +.exercise[ + +- If you're using Play-With-Docker: + ```bash + export DOCKER_HOST=tcp://node3:2375 + ``` + +- Otherwise, use Docker Machine: + ```bash + eval $(docker-machine env node3) + ``` + +] + +--- + +## Checking which node we're talking to + +- Let's use the Docker API to ask "who are you?" to the remote node + +.exercise[ + +- Extract the node name from the output of `docker info`: + ```bash + docker info | grep ^Name + ``` + +] + +This should tell us that we are talking to `node3`. + +Note: it can be useful to use a [custom shell prompt]( +https://github.com/jpetazzo/orchestration-workshop/blob/master/prepare-vms/scripts/postprep.rc#L68) +reflecting the `DOCKER_HOST` variable. + +--- + +## Adding a node through the Docker API + +- We are going to use the same `docker swarm join` command as before + +.exercise[ + +- Add `node3` to the Swarm: + ```bash + docker swarm join --token $TOKEN node1:2377 + ``` + +] + +--- + +## Going back to the local node + +- We need to revert the environment variable(s) that we had set previously + +.exercise[ + +- If you're using Play-With-Docker, just clear `DOCKER_HOST`: + ```bash + unset DOCKER_HOST + ``` + +- Otherwise, use Docker Machine to reset all the relevant variables: + ```bash + eval $(docker-machine env -u) + ``` + +] + +From that point, we are communicating with `node1` again. + +--- + +## Checking the composition of our cluster + +- Now that we're talking to `node1` again, we can use management commands + +.exercise[ + +- Check that the node is here: + ```bash + docker node ls + ``` + +] diff --git a/docs/markmaker.py b/docs/markmaker.py index 08007c46..20550242 100755 --- a/docs/markmaker.py +++ b/docs/markmaker.py @@ -1,6 +1,7 @@ #!/usr/bin/env python # transforms a YAML manifest into a MARKDOWN workshop file +import glob import logging import os import re @@ -24,15 +25,17 @@ def yaml2markdown(inf, outf): logging.debug(titles) toc = gentoc(titles) markdown = markdown.replace("@@TOC@@", toc) + for (s1,s2) in manifest.get("variables", {}).items(): + markdown = markdown.replace(s1, s2) outf.write(markdown) def gentoc(titles, depth=0, chapter=0): if not titles: return "" - if type(titles) == str: + if isinstance(titles, str): return " "*(depth-2) + "- " + titles + "\n" - if type(titles) == list: + if isinstance(titles, list): if depth==0: sep = "\n\n---\n\n" head = "" @@ -56,12 +59,15 @@ def findtitles(markdown): # It returns (epxandedmarkdown,[list of titles]) # The list of titles can be nested. def processchapter(chapter): - if type(chapter) == str: + if isinstance(chapter, unicode): + return processchapter(chapter.encode("utf-8")) + if isinstance(chapter, str): if "\n" in chapter: return (chapter, findtitles(chapter)) if os.path.isfile(chapter): + mdfiles.remove(chapter) return processchapter(open(chapter).read()) - if type(chapter) == list: + if isinstance(chapter, list): chapters = [processchapter(c) for c in chapter] markdown = "\n---\n".join(c[0] for c in chapters) titles = [t for (m,t) in chapters if t] @@ -69,5 +75,6 @@ def processchapter(chapter): raise InvalidChapter(chapter) - +mdfiles = set(glob.glob("*.md")) yaml2markdown(sys.stdin, sys.stdout) +logging.debug("The following files were unused: {}".format(mdfiles)) diff --git a/docs/morenodes.md b/docs/morenodes.md new file mode 100644 index 00000000..848253fe --- /dev/null +++ b/docs/morenodes.md @@ -0,0 +1,236 @@ +## Adding more manager nodes + +- Right now, we have only one manager (node1) + +- If we lose it, we lose quorum - and that's *very bad!* + +- Containers running on other nodes will be fine ... + +- But we won't be able to get or set anything related to the cluster + +- If the manager is permanently gone, we will have to do a manual repair! + +- Nobody wants to do that ... so let's make our cluster highly available + +--- + +class: self-paced + +## Adding more managers + +With Play-With-Docker: + +```bash +TOKEN=$(docker swarm join-token -q manager) +for N in $(seq 4 5); do + export DOCKER_HOST=tcp://node$N:2375 + docker swarm join --token $TOKEN node1:2377 +done +unset DOCKER_HOST +``` + +--- + +class: in-person + +## Building our full cluster + +- We could SSH to nodes 3, 4, 5; and copy-paste the command + +-- + +class: in-person + +- Or we could use the AWESOME POWER OF THE SHELL! + +-- + +class: in-person + +![Mario Red Shell](mario-red-shell.png) + +-- + +class: in-person + +- No, not *that* shell + +--- + +class: in-person + +## Let's form like Swarm-tron + +- Let's get the token, and loop over the remaining nodes with SSH + +.exercise[ + +- Obtain the manager token: + ```bash + TOKEN=$(docker swarm join-token -q manager) + ``` + +- Loop over the 3 remaining nodes: + ```bash + for NODE in node3 node4 node5; do + ssh $NODE docker swarm join --token $TOKEN node1:2377 + done + ``` + +] + +[That was easy.](https://www.youtube.com/watch?v=3YmMNpbFjp0) + +--- + +## You can control the Swarm from any manager node + +.exercise[ + +- Try the following command on a few different nodes: + ```bash + docker node ls + ``` + +] + +On manager nodes: +
you will see the list of nodes, with a `*` denoting +the node you're talking to. + +On non-manager nodes: +
you will get an error message telling you that +the node is not a manager. + +As we saw earlier, you can only control the Swarm through a manager node. + +--- + +class: self-paced + +## Play-With-Docker node status icon + +- If you're using Play-With-Docker, you get node status icons + +- Node status icons are displayed left of the node name + + - No icon = no Swarm mode detected + - Solid blue icon = Swarm manager detected + - Blue outline icon = Swarm worker detected + +![Play-With-Docker icons](pwd-icons.png) + +--- + +## Dynamically changing the role of a node + +- We can change the role of a node on the fly: + + `docker node promote nodeX` → make nodeX a manager +
+ `docker node demote nodeX` → make nodeX a worker + +.exercise[ + +- See the current list of nodes: + ``` + docker node ls + ``` + +- Promote any worker node to be a manager: + ``` + docker node promote + ``` + +] + +--- + +## How many managers do we need? + +- 2N+1 nodes can (and will) tolerate N failures +
(you can have an even number of managers, but there is no point) + +-- + +- 1 manager = no failure + +- 3 managers = 1 failure + +- 5 managers = 2 failures (or 1 failure during 1 maintenance) + +- 7 managers and more = now you might be overdoing it a little bit + +--- + +## Why not have *all* nodes be managers? + +- Intuitively, it's harder to reach consensus in larger groups + +- With Raft, writes have to go to (and be acknowledged by) all nodes + +- More nodes = more network traffic + +- Bigger network = more latency + +--- + +## What would McGyver do? + +- If some of your machines are more than 10ms away from each other, +
+ try to break them down in multiple clusters + (keeping internal latency low) + +- Groups of up to 9 nodes: all of them are managers + +- Groups of 10 nodes and up: pick 5 "stable" nodes to be managers +
+ (Cloud pro-tip: use separate auto-scaling groups for managers and workers) + +- Groups of more than 100 nodes: watch your managers' CPU and RAM + +- Groups of more than 1000 nodes: + + - if you can afford to have fast, stable managers, add more of them + - otherwise, break down your nodes in multiple clusters + +--- + +## What's the upper limit? + +- We don't know! + +- Internal testing at Docker Inc.: 1000-10000 nodes is fine + + - deployed to a single cloud region + + - one of the main take-aways was *"you're gonna need a bigger manager"* + +- Testing by the community: [4700 heterogenous nodes all over the 'net](https://sematext.com/blog/2016/11/14/docker-swarm-lessons-from-swarm3k/) + + - it just works + + - more nodes require more CPU; more containers require more RAM + + - scheduling of large jobs (70000 containers) is slow, though (working on it!) + +--- + +## Real-life deployment methods + +-- + +Running commands manually over SSH + +-- + + (lol jk) + +-- + +- Using your favorite configuration management tool + +- [Docker for AWS](https://docs.docker.com/docker-for-aws/#quickstart) + +- [Docker for Azure](https://docs.docker.com/docker-for-azure/) diff --git a/docs/selfpaced.yml b/docs/selfpaced.yml new file mode 100644 index 00000000..102a9dbc --- /dev/null +++ b/docs/selfpaced.yml @@ -0,0 +1,57 @@ +chapters: +- | + class: title + Docker
Orchestration
Workshop +- intro.md +- | + @@TOC@@ +- - prereqs.md + - versions.md + - | + class: title + + All right! +
+ We're all set. +
+ Let's do this. + - | + name: part-1 + + class: title, self-paced + + Part 1 + - sampleapp.md + - | + class: title + + Scaling out + - swarmkit.md + - creatingswarm.md + - machine.md + - morenodes.md +- - firstservice.md + - ourapponswarm.md +- - operatingswarm.md + - netshoot.md + - swarmnbt.md + - ipsec.md + - updatingservices.md + - healthchecks.md + - nodeinfo.md + - swarmtools.md +- - security.md + - secrets.md + - leastprivilege.md + - namespaces.md + - apiscope.md + - encryptionatrest.md + - logging.md + - metrics.md + - stateful.md + - extratips.md + - end.md +- | + class: title + + Thank you! diff --git a/docs/swarmkit.md b/docs/swarmkit.md index d24a2d48..0aad81ee 100644 --- a/docs/swarmkit.md +++ b/docs/swarmkit.md @@ -7,6 +7,10 @@ - It is a plumbing part of the Docker ecosystem +-- + +.footnote[🐳 Did you know that кит means "whale" in Russian?] + --- ## SwarmKit features @@ -145,853 +149,3 @@ You will get an error message: ``` Error response from daemon: This node is not a swarm manager. [...] ``` - ---- - -# Creating our first Swarm - -- The cluster is initialized with `docker swarm init` - -- This should be executed on a first, seed node - -- .warning[DO NOT execute `docker swarm init` on multiple nodes!] - - You would have multiple disjoint clusters. - -.exercise[ - -- Create our cluster from node1: - ```bash - docker swarm init - ``` - -] - --- - -class: advertise-addr - -If Docker tells you that it `could not choose an IP address to advertise`, see next slide! - ---- - -class: advertise-addr - -## IP address to advertise - -- When running in Swarm mode, each node *advertises* its address to the others -
- (i.e. it tells them *"you can contact me on 10.1.2.3:2377"*) - -- If the node has only one IP address (other than 127.0.0.1), it is used automatically - -- If the node has multiple IP addresses, you **must** specify which one to use -
- (Docker refuses to pick one randomly) - -- You can specify an IP address or an interface name -
(in the latter case, Docker will read the IP address of the interface and use it) - -- You can also specify a port number -
(otherwise, the default port 2377 will be used) - ---- - -class: advertise-addr - -## Which IP address should be advertised? - -- If your nodes have only one IP address, it's safe to let autodetection do the job - - .small[(Except if your instances have different private and public addresses, e.g. - on EC2, and you are building a Swarm involving nodes inside and outside the - private network: then you should advertise the public address.)] - -- If your nodes have multiple IP addresses, pick an address which is reachable - *by every other node* of the Swarm - -- If you are using [play-with-docker](http://play-with-docker.com/), use the IP - address shown next to the node name - - .small[(This is the address of your node on your private internal overlay network. - The other address that you might see is the address of your node on the - `docker_gwbridge` network, which is used for outbound traffic.)] - -Examples: - -```bash -docker swarm init --advertise-addr 10.0.9.2 -docker swarm init --advertise-addr eth0:7777 -``` - ---- - -class: extra-details - -## Using a separate interface for the data path - -- You can use different interfaces (or IP addresses) for control and data - -- You set the _control plane path_ with `--advertise-addr` - - (This will be used for SwarmKit manager/worker communication, leader election, etc.) - -- You set the _data plane path_ with `--data-path-addr` - - (This will be used for traffic between containers) - -- Both flags can accept either an IP address, or an interface name - - (When specifying an interface name, Docker will use its first IP address) - ---- - -## Token generation - -- In the output of `docker swarm init`, we have a message - confirming that our node is now the (single) manager: - - ``` - Swarm initialized: current node (8jud...) is now a manager. - ``` - -- Docker generated two security tokens (like passphrases or passwords) for our cluster - -- The CLI shows us the command to use on other nodes to add them to the cluster using the "worker" - security token: - - ``` - To add a worker to this swarm, run the following command: - docker swarm join \ - --token SWMTKN-1-59fl4ak4nqjmao1ofttrc4eprhrola2l87... \ - 172.31.4.182:2377 - ``` - ---- - -class: extra-details - -## Checking that Swarm mode is enabled - -.exercise[ - -- Run the traditional `docker info` command: - ```bash - docker info - ``` - -] - -The output should include: - -``` -Swarm: active - NodeID: 8jud7o8dax3zxbags3f8yox4b - Is Manager: true - ClusterID: 2vcw2oa9rjps3a24m91xhvv0c - ... -``` - ---- - -## Running our first Swarm mode command - -- Let's retry the exact same command as earlier - -.exercise[ - -- List the nodes (well, the only node) of our cluster: - ```bash - docker node ls - ``` - -] - -The output should look like the following: -``` -ID HOSTNAME STATUS AVAILABILITY MANAGER STATUS -8jud...ox4b * node1 Ready Active Leader -``` - ---- - -## Adding nodes to the Swarm - -- A cluster with one node is not a lot of fun - -- Let's add `node2`! - -- We need the token that was shown earlier - --- - -- You wrote it down, right? - --- - -- Don't panic, we can easily see it again 😏 - ---- - -## Adding nodes to the Swarm - -.exercise[ - -- Show the token again: - ```bash - docker swarm join-token worker - ``` - -- Switch to `node2` - -- Copy-paste the `docker swarm join ...` command -
(that was displayed just before) - -] - ---- - -class: extra-details - -## Check that the node was added correctly - -- Stay on `node2` for now! - -.exercise[ - -- We can still use `docker info` to verify that the node is part of the Swarm: - ```bash - docker info | grep ^Swarm - ``` - -] - -- However, Swarm commands will not work; try, for instance: - ``` - docker node ls - ``` - -- This is because the node that we added is currently a *worker* - -- Only *managers* can accept Swarm-specific commands - ---- - -## View our two-node cluster - -- Let's go back to `node1` and see what our cluster looks like - -.exercise[ - -- Switch back to `node1` - -- View the cluster from `node1`, which is a manager: - ```bash - docker node ls - ``` - -] - -The output should be similar to the following: -``` -ID HOSTNAME STATUS AVAILABILITY MANAGER STATUS -8jud...ox4b * node1 Ready Active Leader -ehb0...4fvx node2 Ready Active -``` - ---- - -class: docker-machine - -## Adding nodes using the Docker API - -- We don't have to SSH into the other nodes, we can use the Docker API - -- If you are using Play-With-Docker: - - - the nodes expose the Docker API over port 2375/tcp, without authentication - - - we will connect by setting the `DOCKER_HOST` environment variable - -- Otherwise: - - - the nodes expose the Docker API over port 2376/tcp, with TLS mutual authentication - - - we will use Docker Machine to set the correct environment variables -
(the nodes have been suitably pre-configured to be controlled through `node1`) - ---- - -class: docker-machine - -# Docker Machine - -- Docker Machine has two primary uses: - - - provisioning cloud instances running the Docker Engine - - - managing local Docker VMs within e.g. VirtualBox - -- Docker Machine is purely optional - -- It makes it easy to create, upgrade, manage... Docker hosts: - - - on your favorite cloud provider - - - locally (e.g. to test clustering, or different versions) - - - across different cloud providers - ---- - -class: self-paced, docker-machine - -## If you're using Play-With-Docker ... - -- You won't need to use Docker Machine - -- Instead, to "talk" to another node, we'll just set `DOCKER_HOST` - -- You can skip the exercises telling you to do things with Docker Machine! - ---- - -class: docker-machine - -## Docker Machine basic usage - -- We will learn two commands: - - - `docker-machine ls` (list existing hosts) - - - `docker-machine env` (switch to a specific host) - -.exercise[ - -- List configured hosts: - ```bash - docker-machine ls - ``` - -] - -You should see your 5 nodes. - ---- - -class: in-person, docker-machine - -## How did we make our 5 nodes show up there? - -*For the curious...* - -- This was done by our VM provisioning scripts - -- After setting up everything else, `node1` adds the 5 nodes - to the local Docker Machine configuration - (located in `$HOME/.docker/machine`) - -- Nodes are added using [Docker Machine generic driver](https://docs.docker.com/machine/drivers/generic/) - - (It skips machine provisioning and jumps straight to the configuration phase) - -- Docker Machine creates TLS certificates and deploys them to the nodes through SSH - ---- - -class: docker-machine - -## Using Docker Machine to communicate with a node - -- To select a node, use `eval $(docker-machine env nodeX)` - -- This sets a number of environment variables - -- To unset these variables, use `eval $(docker-machine env -u)` - -.exercise[ - -- View the variables used by Docker Machine: - ```bash - docker-machine env node3 - ``` - -] - -(This shows which variables *would* be set by Docker Machine; but it doesn't change them.) - ---- - -class: docker-machine - -## Getting the token - -- First, let's store the join token in a variable - -- This must be done from a manager - -.exercise[ - -- Make sure we talk to the local node, or `node1`: - ```bash - eval $(docker-machine env -u) - ``` - -- Get the join token: - ```bash - TOKEN=$(docker swarm join-token -q worker) - ``` - -] - ---- - -class: docker-machine - -## Change the node targeted by the Docker CLI - -- We need to set the right environment variables to communicate with `node3` - -.exercise[ - -- If you're using Play-With-Docker: - ```bash - export DOCKER_HOST=tcp://node3:2375 - ``` - -- Otherwise, use Docker Machine: - ```bash - eval $(docker-machine env node3) - ``` - -] - ---- - -class: docker-machine - -## Checking which node we're talking to - -- Let's use the Docker API to ask "who are you?" to the remote node - -.exercise[ - -- Extract the node name from the output of `docker info`: - ```bash - docker info | grep ^Name - ``` - -] - -This should tell us that we are talking to `node3`. - -Note: it can be useful to use a [custom shell prompt]( -https://github.com/jpetazzo/orchestration-workshop/blob/master/prepare-vms/scripts/postprep.rc#L68) -reflecting the `DOCKER_HOST` variable. - ---- - -class: docker-machine - -## Adding a node through the Docker API - -- We are going to use the same `docker swarm join` command as before - -.exercise[ - -- Add `node3` to the Swarm: - ```bash - docker swarm join --token $TOKEN node1:2377 - ``` - -] - ---- - -class: docker-machine - -## Going back to the local node - -- We need to revert the environment variable(s) that we had set previously - -.exercise[ - -- If you're using Play-With-Docker, just clear `DOCKER_HOST`: - ```bash - unset DOCKER_HOST - ``` - -- Otherwise, use Docker Machine to reset all the relevant variables: - ```bash - eval $(docker-machine env -u) - ``` - -] - -From that point, we are communicating with `node1` again. - ---- - -class: docker-machine - -## Checking the composition of our cluster - -- Now that we're talking to `node1` again, we can use management commands - -.exercise[ - -- Check that the node is here: - ```bash - docker node ls - ``` - -] - ---- - -class: under-the-hood - -## Under the hood: docker swarm init - -When we do `docker swarm init`: - -- a keypair is created for the root CA of our Swarm - -- a keypair is created for the first node - -- a certificate is issued for this node - -- the join tokens are created - ---- - -class: under-the-hood - -## Under the hood: join tokens - -There is one token to *join as a worker*, and another to *join as a manager*. - -The join tokens have two parts: - -- a secret key (preventing unauthorized nodes from joining) - -- a fingerprint of the root CA certificate (preventing MITM attacks) - -If a token is compromised, it can be rotated instantly with: -``` -docker swarm join-token --rotate -``` - ---- - -class: under-the-hood - -## Under the hood: docker swarm join - -When a node joins the Swarm: - -- it is issued its own keypair, signed by the root CA - -- if the node is a manager: - - - it joins the Raft consensus - - it connects to the current leader - - it accepts connections from worker nodes - -- if the node is a worker: - - - it connects to one of the managers (leader or follower) - ---- - -class: under-the-hood - -## Under the hood: cluster communication - -- The *control plane* is encrypted with AES-GCM; keys are rotated every 12 hours - -- Authentication is done with mutual TLS; certificates are rotated every 90 days - - (`docker swarm update` allows to change this delay or to use an external CA) - -- The *data plane* (communication between containers) is not encrypted by default - - (but this can be activated on a by-network basis, using IPSEC, - leveraging hardware crypto if available) - ---- - -class: under-the-hood - -## Under the hood: I want to know more! - -Revisit SwarmKit concepts: - -- Docker 1.12 Swarm Mode Deep Dive Part 1: Topology - ([video](https://www.youtube.com/watch?v=dooPhkXT9yI)) - -- Docker 1.12 Swarm Mode Deep Dive Part 2: Orchestration - ([video](https://www.youtube.com/watch?v=_F6PSP-qhdA)) - -Some presentations from the Docker Distributed Systems Summit in Berlin: - -- Heart of the SwarmKit: Topology Management - ([slides](https://speakerdeck.com/aluzzardi/heart-of-the-swarmkit-topology-management)) - -- Heart of the SwarmKit: Store, Topology & Object Model - ([slides](http://www.slideshare.net/Docker/heart-of-the-swarmkit-store-topology-object-model)) - ([video](https://www.youtube.com/watch?v=EmePhjGnCXY)) - ---- - -## Adding more manager nodes - -- Right now, we have only one manager (node1) - -- If we lose it, we lose quorum - and that's *very bad!* - -- Containers running on other nodes will be fine ... - -- But we won't be able to get or set anything related to the cluster - -- If the manager is permanently gone, we will have to do a manual repair! - -- Nobody wants to do that ... so let's make our cluster highly available - ---- - -class: self-paced - -## Adding more managers - -With Play-With-Docker: - -```bash -TOKEN=$(docker swarm join-token -q manager) -for N in $(seq 4 5); do - export DOCKER_HOST=tcp://node$N:2375 - docker swarm join --token $TOKEN node1:2377 -done -unset DOCKER_HOST -``` - ---- - -class: docker-machine - -## Adding more managers - -With Docker Machine: - -```bash -TOKEN=$(docker swarm join-token -q manager) -for N in $(seq 4 5); do - eval $(docker-machine env node$N) - docker swarm join --token $TOKEN node1:2377 -done -eval $(docker-machine env -u) -``` - ---- - -class: in-person - -## Building our full cluster - -- We could SSH to nodes 3, 4, 5; and copy-paste the command - --- - -class: in-person - -- Or we could use the AWESOME POWER OF THE SHELL! - --- - -class: in-person - -![Mario Red Shell](mario-red-shell.png) - --- - -class: in-person - -- No, not *that* shell - ---- - -class: in-person - -## Let's form like Swarm-tron - -- Let's get the token, and loop over the remaining nodes with SSH - -.exercise[ - -- Obtain the manager token: - ```bash - TOKEN=$(docker swarm join-token -q manager) - ``` - -- Loop over the 3 remaining nodes: - ```bash - for NODE in node3 node4 node5; do - ssh $NODE docker swarm join --token $TOKEN node1:2377 - done - ``` - -] - -[That was easy.](https://www.youtube.com/watch?v=3YmMNpbFjp0) - ---- - -## You can control the Swarm from any manager node - -.exercise[ - -- Try the following command on a few different nodes: - ```bash - docker node ls - ``` - -] - -On manager nodes: -
you will see the list of nodes, with a `*` denoting -the node you're talking to. - -On non-manager nodes: -
you will get an error message telling you that -the node is not a manager. - -As we saw earlier, you can only control the Swarm through a manager node. - ---- - -class: self-paced - -## Play-With-Docker node status icon - -- If you're using Play-With-Docker, you get node status icons - -- Node status icons are displayed left of the node name - - - No icon = no Swarm mode detected - - Solid blue icon = Swarm manager detected - - Blue outline icon = Swarm worker detected - -![Play-With-Docker icons](pwd-icons.png) - ---- - -## Dynamically changing the role of a node - -- We can change the role of a node on the fly: - - `docker node promote XXX` → make XXX a manager -
- `docker node demote XXX` → make XXX a worker - -.exercise[ - -- See the current list of nodes: - ``` - docker node ls - ``` - -- Promote any worker node to be a manager: - ``` - docker node promote - ``` - -] - ---- - -## How many managers do we need? - -- 2N+1 nodes can (and will) tolerate N failures -
(you can have an even number of managers, but there is no point) - --- - -- 1 manager = no failure - -- 3 managers = 1 failure - -- 5 managers = 2 failures (or 1 failure during 1 maintenance) - -- 7 managers and more = now you might be overdoing it a little bit - ---- - -## Why not have *all* nodes be managers? - -- Intuitively, it's harder to reach consensus in larger groups - -- With Raft, writes have to go to (and be acknowledged by) all nodes - -- More nodes = more network traffic - -- Bigger network = more latency - ---- - -## What would McGyver do? - -- If some of your machines are more than 10ms away from each other, -
- try to break them down in multiple clusters - (keeping internal latency low) - -- Groups of up to 9 nodes: all of them are managers - -- Groups of 10 nodes and up: pick 5 "stable" nodes to be managers -
- (Cloud pro-tip: use separate auto-scaling groups for managers and workers) - -- Groups of more than 100 nodes: watch your managers' CPU and RAM - -- Groups of more than 1000 nodes: - - - if you can afford to have fast, stable managers, add more of them - - otherwise, break down your nodes in multiple clusters - ---- - -## What's the upper limit? - -- We don't know! - -- Internal testing at Docker Inc.: 1000-10000 nodes is fine - - - deployed to a single cloud region - - - one of the main take-aways was *"you're gonna need a bigger manager"* - -- Testing by the community: [4700 heterogenous nodes all over the 'net](https://sematext.com/blog/2016/11/14/docker-swarm-lessons-from-swarm3k/) - - - it just works - - - more nodes require more CPU; more containers require more RAM - - - scheduling of large jobs (70000 containers) is slow, though (working on it!) - ---- - -## Real-life deployment methods - --- Running commands manually over SSH - --- - - (lol jk) - --- - -- Using your favorite configuration management tool - -- [Docker for AWS](https://docs.docker.com/docker-for-aws/#quickstart) - -- [Docker for Azure](https://docs.docker.com/docker-for-azure/) diff --git a/docs/workshop.html b/docs/workshop.html index 0a5c4f91..400d268d 100644 --- a/docs/workshop.html +++ b/docs/workshop.html @@ -7,7 +7,8 @@ -
+
+

Loading ...

The slides should show up here. If they don't, it might be because you are accessing this file directly from your filesystem. It needs to be served from a web server. You can try this: @@ -26,7 +27,7 @@ sourceUrl: 'workshop.md', ratio: '16:9', highlightSpans: true, - excludedClasses: ["extra-details", "self-paced"] + excludedClasses: ["self-paced", "snap"] });