ceph-ansible

Commit Graph

Author	SHA1	Message	Date
Dimitri Savineau	c97f90af92	ceph-facts: move device facts to its own file Instead of reusing the condition 'inventory_hostname in groups[osds]' on each device facts tasks then we can move all the tasks into a dedicated file and set the condition on the import_tasks statement. Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `d704b05e52`)	2021-07-02 22:21:32 +02:00
Dimitri Savineau	efd80b07de	ceph-validate: check logical volumes We currently don't check if the logical volume used in lvm_volumes list for either bluestore data/db/wal or filestore data/journal exist. We're only doing this on raw devices for batch scenario. Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `55bca07cb6`)	2021-07-02 22:21:32 +02:00
Dimitri Savineau	b61b20f25b	ceph-validate: check db/journal/wal devices too When using dedicated devices for db/journal/wal objecstore with ceph-volume lvm batch then we should also validate that those devices exist and don't use a gpt partition table in addition of the devices and lvm_volume.data variables. Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `808e7106de`)	2021-07-02 22:21:32 +02:00
Dimitri Savineau	71c0f2d6e1	ceph-validate: use root device from ansible_mounts Instead of using findmnt command to find the device associated to the root mount point then we can use the ansible_mounts fact. Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `7e50380f7f`)	2021-07-02 22:21:32 +02:00
Dimitri Savineau	5476d606a3	ceph-validate: do not resolve devices This is already done in the ceph-facts role. Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `0df99dda8d`)	2021-07-02 22:21:32 +02:00
Dimitri Savineau	11bb1ece2e	ceph-validate: check block presence first Instead of doing two parted calls we can check first if the device exist and then test the partition table. Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `14d458b3b4`)	2021-07-02 22:21:32 +02:00
Dimitri Savineau	a3b5b15160	ceph-validate: check devices from lvm_volumes `2888c08` introduced a regression as the check_devices tasks file was only included based on the devices variable. But that file also validate some devices from the lvm_volumes variable. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1906022 Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `ac0342b72e`)	2021-07-02 22:21:32 +02:00
Dimitri Savineau	6bda641584	prometheus: fix prometheus target url The prometheus service isn't binding on localhost. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1933560 Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `1d56818658`)	2021-07-02 14:37:49 -04:00
Dimitri Savineau	4ce49927f7	container: set tcmalloc value by default All ceph daemons need to have the TCMALLOC_MAX_TOTAL_THREAD_CACHE_BYTES environment variable set to 128MB by default in container setup. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1970913 Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `9758e3c513`)	2021-07-01 15:46:06 +02:00
Guillaume Abrioux	2957b69c8b	ceph_key: handle error in a better way When calling the `ceph_key` module with `state: info`, if the ceph command called fails, the actual error is hidden by the module which makes it pretty difficult to troubleshoot. The current code always states that if rc is not equal to 0 the keyring doesn't exist. `state: info` should always return the actual rc, stdout and stderr. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1964889 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `d58500ade0`)	2021-06-30 20:34:30 +02:00
Boris Ranto	a348116bf7	dashboard: Add new prometheus alert It was requested for us to update our alerting definitions to include a slow OSD Ops health check. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1951664 Signed-off-by: Boris Ranto <branto@redhat.com> (cherry picked from commit `2491d4e004`)	2021-06-30 15:30:44 +02:00
VasishtaShastry	5a31701229	Container: Fixing service name lvm2-lvmetad Playbook failing saying: msg: 'Could not find the requested service lvmetad: host' Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1955040 Signed-off-by: VasishtaShastry <vipin.indiasmg@gmail.com> (cherry picked from commit `e49c38f8b7`)	2021-06-17 08:14:37 +02:00
Guillaume Abrioux	cb4dbd6916	multisite: fix bug during switch2containers When running the switch-to-containers playbook with multisite enabled, the fact "rgw_instances" is only set for the node being processed (serial: 1), the consequence of that is that the set_fact of 'rgw_instances_all' can't iterate over all rgw node in order to look up each 'rgw_instances_host'. Adding a condition checking whether hostvars[item]["rgw_instances_host"] is defined fixes this issue. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1967926 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `8279d14d32`)	2021-06-17 07:20:12 +02:00
Guillaume Abrioux	69cc3b82c0	nfs: do no copy client.bootstrap-rgw when using mds There's no need to copy this keyring when using nfs with mds Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `8dbee99882`)	2021-06-16 09:33:59 +02:00
Guillaume Abrioux	82b06afc96	rolling_update: fix mon+rgw/multisite collocation When monitors and rgw are collocated with multisite enabled, the rolling_update playbook fails because during the workflow, we run some radosgw-admin commands very early on the first mon even though this is the monitor being upgraded, it means the container doesn't exist since it was stopped. This block is relevant only for scaling out rgw daemons or initial deployment. In rolling_update workflow, it is not needed so let's skip it. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1970232 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `f7166cccbf`)	2021-06-14 09:27:13 +02:00
Guillaume Abrioux	69a80aea9e	dashboard: set cookie_secure in grafana When using grafana behind https `cookie_secure` should be set to `true`. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1966880 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `4daed1f137`)	2021-06-07 15:12:33 +02:00
Guillaume Abrioux	ac0a5c1e68	dashboard: fix typo introduced during backport during backport of `c8b92deba1` the pattern should have been s/monitoring_group_name/grafana_server_group_name/ Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1964907 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2021-05-26 15:23:01 +02:00
Guillaume Abrioux	10ed26f14d	crash: fix --limit deployments (containers) ceph-crash deployments is broken when ceph-ansible playbook is called with --limit in containerized contexts since we don't set `container_exec_cmd` on the first monitor. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1964835 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2021-05-26 12:58:16 +02:00
Guillaume Abrioux	e63a487a88	prometheus: enforce osd nodes in templates When osd nodes are collocated in the clients group (HCI context for instance), the current logic will exclude osd nodes since they are present in the client group. The best fix would be to exclude clients node only when they are not member of another group but for now, as a workaround, we can enforce the addition of osd nodes to fix this specific case. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1947695 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `664dae0564`)	2021-05-26 10:33:19 +02:00
Guillaume Abrioux	71bd606855	container: conditionnally disable lvmetad Enabling lvmetad in containerized deployments on el7 based OS might cause issues. This commit make it possible to disable this service if needed. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1955040 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2021-05-25 16:50:45 +02:00
Brad Hubbard	b6d75bee2e	Make sure the repo url contains the correct arch We can end up with an arm only repo unless we are specific about the architecture we require. Brings the deb code in line with the rpm equivalent. Signed-off-by: Brad Hubbard <bhubbard@redhat.com> (cherry picked from commit `267cce9e83`)	2021-05-17 08:56:50 +02:00
Dimitri Savineau	3a9a7bcf73	ceph-rgw: fix pg_autoscale_mode for pool The pg_autoscale_mode for rgw pools introduced in `9f03a52` was wrong and was missing a `value` keyword because `rgw_create_pools` is a dict. Fixes: #6516 Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `a670982a38`)	2021-05-10 21:11:40 +02:00
Guillaume Abrioux	f185e409da	nfs: get org.ganesha.nfsd.conf from container Since we need to revert `33bfb10`, this is an alternative to initial approach. We can avoid maintaining this file since it is present in container image. The idea is to simply get it from the image container and write it to the host. Fixes: #6501 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `e6d8b058ba`)	2021-05-07 16:34:46 +02:00
Guillaume Abrioux	bf570600a2	nfs: remove legacy task This fact is never used, let's remove the task. Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `0772b3d28d`)	2021-05-06 10:15:31 +02:00
Guillaume Abrioux	1ac5e36096	nfs: rename two tasks set the name of those tasks accordingly with the fact name being set. Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `d3d3d01528`)	2021-05-06 10:15:31 +02:00
Guillaume Abrioux	c79b866435	validate: check virtual_ips variable This commit checks the length of `virtual_ips` doesn't exceed the length of `groups[rgwloadbalancer_group_name]`. It also ensure this variable is defined when `groups[rgwloadbalancer_group_name]` contains at least one node. Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `ae196bf946`)	2021-05-05 09:55:33 +02:00
Benoît Knecht	6451d3681b	ceph-rgw-loadbalancer: Fix keepalived master selection While `2ca33641` fixed a bug in the way the `keepalived.conf.j2` template matched hostnames to set the VRRP `MASTER`/`BACKUP` states, it also introduced a regression in the case where `virtual_ips` is a list of more than one IP address. The previous behavior would result in each host in the `rgwloadbalancers` group to be `MASTER` for one of the `virtual_ips`, but the new behavior caused the first host to be `MASTER` for all the IP address in `virtual_ips`. This commit restores the original behavior. Signed-off-by: Benoît Knecht <bknecht@protonmail.ch> (cherry picked from commit `3116f46422`)	2021-05-05 09:55:33 +02:00
Seena Fallah	36dc972e09	ceph-osd: allow to use ceph_tcmalloc_max_total_thread_cache for bluestore TCMALLOC_MAX_TOTAL_THREAD_CACHE_BYTES is for both bluestore and filestore Signed-off-by: Seena Fallah <seenafallah@gmail.com> (cherry picked from commit `41295f0ef6`)	2021-04-29 07:34:43 +02:00
Benoît Knecht	80cf5b731b	ceph-mon: Fix check mode for deploy monitor tasks Skip the `get initial keyring when it already exists` task when both commands whose `stdout` output it requires have been skipped (e.g. when running in check mode). Signed-off-by: Benoît Knecht <bknecht@protonmail.ch> (cherry picked from commit `2437f14581`)	2021-04-28 15:18:49 +02:00
Francesco Pantano	693a8087bc	Config the monitoring stack components api urls using a VIP When dashboard_frontend_vip is provided, all the services should be configured using the related VIP. A new VIP variable is added for both prometheus and alertmanager: we're already able to properly config the grafana vip using dashboard_frontend_vip variable. This change adds the same variable for both prometheus and alertmanager. Signed-off-by: Francesco Pantano <fpantano@redhat.com> (cherry picked from commit `441651638d`)	2021-04-28 08:54:09 +02:00
Benoît Knecht	fb35ca364b	ceph-rgw-loadbalancer: Fix rgw_ports fact The `set_fact rgw_ports` task was failing due to a templating error, because `hostvars[item].rgw_instances` is a list, but it was treated as if it was a dictionary. Another issue was the fact that the `unique` filter only applied to the list being appended to `rgw_ports` instead of the entire list, which means it was possible to have duplicate items. Lastly, `rgw_ports` would have been a list of integers, but the `seport` module expects a list of strings. This commit fixes all of the issues above, allowing the `ceph-rgw-loadbalancer` role to work on systems with SELinux enabled. Signed-off-by: Benoît Knecht <bknecht@protonmail.ch> (cherry picked from commit `c078513475`)	2021-04-15 13:20:55 +02:00
Guillaume Abrioux	8f5abc6d3e	container/systemd: ensure /var/log/ceph exists This adds a `ExecStartPre=-/usr/bin/mkdir -p /var/log/ceph` in all systemd service templates for all ceph daemon. This is specific to RHCS after a Leapp upgrade is done. Indeed, the `/var/log/ceph` seems to be removed after the upgrade. In order to work around this issue let's ensure the directory is present before trying to start the containers with podman. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1949489 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `bab403b603`)	2021-04-14 20:05:24 +02:00
Guillaume Abrioux	c8d7994117	rbdmirror: add retries/until when configuring mirroring `configure_mirroring.yml` is called right after the daemon is started. Sometimes, it can happen the first task in `configure_mirroring.yml` is run while the daemon isn't yet ready, adding a retries/until on that task should help to avoid causing the playbook to fail. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1944996 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `b1e7e1ad0f`)	2021-04-14 16:12:49 +02:00
Guillaume Abrioux	568d1d6427	docker2podman: skip some role imports from handler when running docker-to-podman playbook, there's no need to call `ceph-config` and `ceph-rgw` from the role `ceph-handler`. It can even have side effects when coming from a baremetal cluster that was previously migrated using the switch-to-containers playbook. Indeed it might complain about missing .target systemd unit since they are removed during that migration. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1944999 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `70f19be367`)	2021-04-12 13:30:17 +02:00
Guillaume Abrioux	ae452a86dc	common: selinux tasks related refactor This moves some task from the `ceph-nfs` role in `ceph-common` since some of them are needed in `ceph-rgwloadbalancer` role. This avoids duplicated tasks. Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `d0442d81b9`)	2021-04-06 15:08:51 +02:00
Guillaume Abrioux	b02c5e8db7	rgw-loadbalancers: add all rgw_ports to http_port_t type This adds all rgw ports to the http_port_t selinux type so it allows haproxy to connect to those ports in order to avoid AVC. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1923890 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `6bbb90198b`)	2021-04-06 15:08:51 +02:00
kalebskeithley	e63e3a65b4	rgw-loadbalancer: Update haproxy.cfg.j2 haproxy gets an AVC when configured to connect to port 8081 This commit adds a snippet regarding haproxy in a selinux environment Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1923890 Signed-off-by: Kaleb S KEITHLEY <kkeithle@redhat.com> (cherry picked from commit `9e7f22a071`)	2021-04-06 15:08:51 +02:00
Dimitri Savineau	501e33bc6a	container/registry: use password from stdin Pass the password variable via stdin for the registry login authentication. This allows to remove the no_log statement and see the task output without displaying the password value. Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `a0e1a450d3`)	2021-04-02 13:14:08 +02:00
Guillaume Abrioux	b7a699f75d	rgw: supports pg_autoscale_mode option for pool creation Support enabling/disabling the pg autoscaler for rgw pools. Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `9f03a527ba`)	2021-04-01 15:33:09 +02:00
Guillaume Abrioux	15a0591615	dashboard: support prometheus storage.tsdb.retention.time parameter This commit adds the parameter `--storage.tsdb.retention.time` to the prometheus systemd unit template. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1928000 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `b60c61ce45`)	2021-04-01 14:52:50 +02:00
Guillaume Abrioux	c8afa768ab	nfs: set idmap config for Ceph-NFS Currently NFS Ganesha (ceph-nfs) consumes /etc/idmapd.conf, which controls mapping of user/owner identities under NFSv4+. With containerized service deployment, this file is an immutable part of the container image and cannot be modified. Here we provide group variables, and a taskk and templates for the ceph-nfs role, to set the path of the idmap configuration file and to make the most common adjustment to the contents of that file -- namely to set the 'Domain'. We default the path to /etc/ganesha/idmap.conf so that we will not conflict with /etc/idmapd.conf on the controller nodes where ganesha runs. NFSv4 clients, as used for example by the Cinder NFS driver, consume /etc/idmapd.conf and may require different settings than what is wanted for NFS Ganesha. Additionally, because we already bind /etc/ganesha from the host into the ceph-nfs container, the file NFS Ganesha consumes will no longer be an immutable part of the container. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1925646 Signed-off-by: Tom Barron tpb@dyncloud.net Co-Authored-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `2db2208e40`)	2021-04-01 14:52:25 +02:00
Guillaume Abrioux	93fd5532ba	defaults: add a comment about `igw_network` This add a quick documentation in ceph-defaults about `igw_network` Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `c5728bdc63`)	2021-03-29 11:24:14 +02:00
Guillaume Abrioux	52a0b222c1	dashboard: support igw nodes with dedicated subnet This adds the possibility to deploy the dashboard with igw nodes using a dedicated subnet. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1926170 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `c33de174f1`)	2021-03-26 21:25:57 +01:00
VasishtaShastry	af6abb7125	Peer addition won't be skipped if remote is not in peer rbd-mirroring is not configured as adding peer is getting skipped. Peer addition should not get skipped if its not added already Closes - https://bugzilla.redhat.com/show_bug.cgi?id=1942444 Signed-off-by: VasishtaShastry <vipin.indiasmg@gmail.com> (cherry picked from commit `006998e804`)	2021-03-26 19:14:49 +01:00
Guillaume Abrioux	50b95baa32	convert some missed `ansible_`` calls to `ansible_facts['']` This converts some missed calls to `ansible_*` that were missed in initial PR #6312 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `0163ecc924`)	2021-03-26 00:05:33 +01:00
Guillaume Abrioux	d6fcd78e72	clients: build filtered clients group early when the group `_filtered_clients` is built, the order can change from the original `clients` group which can cause issues since we run `ceph-container-engine` on the first client only. It means later in the playbook we can make call to the container CLI on a node where the container engine wasn't installed. Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `a112572734`)	2021-03-26 00:05:33 +01:00
Alex Schultz	7c6783acb1	Disable facts by default in ansible.cfg As a continuation of `a7f2fa73e6`, this change switches fact injection to off by default in the provided ansible.cfg. Signed-off-by: Alex Schultz <aschultz@redhat.com> (cherry picked from commit `db031a4993`) (cherry picked from commit `5fa4ff5ed3`)	2021-03-26 00:05:33 +01:00
Alex Schultz	815ea7765f	Use ansible_facts It has come to our attention that using ansible_* vars that are populated with INJECT_FACTS_AS_VARS=True is not very performant. In order to be able to support setting that to off, we need to update the references to use ansible_facts[<thing>] instead of ansible_<thing>. Related: ansible#73654 Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1935406 Signed-off-by: Alex Schultz <aschultz@redhat.com> (cherry picked from commit `a7f2fa73e6`)	2021-03-26 00:05:33 +01:00
Guillaume Abrioux	1fe44154de	iscsi: fetch right repo from shaman due to recent changes in shaman, we must fetch the right repo by filtering on the desired architecture. Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `5801171b37`)	2021-03-24 09:20:24 +01:00
Guillaume Abrioux	802705ff9b	facts: fix nfs/external cluster scenario These tasks shouldn't be run when at least 1 monitor isn't present in the inventory. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1937997 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `ccd1cbb732`)	2021-03-18 06:40:44 +01:00
Guillaume Abrioux	8e30a3c9f8	config: reset num_osds When collocating OSDs with other daemon, `num_osds` is incorrectly calculated because `ceph-config` is called multiple times. Indeed, the following code: ``` num_osds: "{{ lvm_list.stdout \| default('{}') \| from_json \| length \| int + num_osds \| default(0) \| int }}" ``` makes `num_osds` be incremented each time `ceph-config` is called. We have to reset it in order to get the correct number of expected OSDs. Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `31a0f2653d`)	2021-03-17 17:35:37 +01:00
Matthew Vernon	d449d15d4d	ceph-osd: add prepare_osd tag to lvm-batch scenario Sometimes it's useful to be able to skip the OSD creation step when running ceph-ansible (cf #1777). The lvm scenario has a prepare_osd tag on the relevant play. This commit adds the same tag to the lvm-batch scenario. Signed-off-by: Matthew Vernon <mv3@sanger.ac.uk> (cherry picked from commit `88d119e95a`)	2021-03-12 15:44:32 +01:00
Guillaume Abrioux	8b69451652	dashboard: add missing parameter in `ceph_cmd` the `ceph_cmd` fact is missing the `--net=host` parameter. Some tasks consuming this fact can fail like following: ``` Error: error configuring network namespace for container b8ec913db1fb694ae683faf202680de7a59c714a004e533aba87e8503d29261f: Missing CNI default network ``` Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1931365 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `f143b1a647`)	2021-03-12 09:38:46 +01:00
Guillaume Abrioux	32ad0f6fe7	common: ensure shaman returns right repo Due to recent changes in shaman, there's a chance it returns the wrong repository from architecture point of view. We can query shaman and ask for the correct architecture to get around this. Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `39649f0ce8`)	2021-03-10 17:17:33 -05:00
Dimitri Savineau	09d6706697	debian/uca: remove the handler notification The "update apt cache" in the ceph-handler role was never called and the handler trigger after adding the uca repository doesn't exist at all. Instead of using a handler for that we can just set the update_cache parameter to true like the other apt_repository tasks. Resolve merge conflict from cherry-picking this commit. Signed-off-by: Dimitri Savineau <dsavinea@redhat.com>	2021-03-03 14:50:45 +01:00
Matthew Vernon	42b571b11f	Fix typo and broken link for documenting RGW frontends http://docs.ceph.com/docs/nautilus/radosgw/frontends/ 404s so replace it with a working "latest" docs link, and correct the spelling of "additional" while I'm at it. Signed-off-by: Matthew Vernon <mv3@sanger.ac.uk> (cherry picked from commit `847611048e`)	2021-03-03 14:19:52 +01:00
Dimitri Savineau	5572c907ee	ceph-common: enable rhcs tools repo for monitoring The monitoring node running grafana needs the rhcs tools repostory enabled in non containerized deployment to be able to install the ceph-grafana-dashboards rpm package. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1918650 Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `e4dd0067c6`)	2021-02-10 09:58:03 +01:00
Guillaume Abrioux	dd204d9e2f	rgw: fix a typo in multisite if `rgw_zonegroupmaster` is not defined at the rgw instance level in `rgw_instances` it will fallback to a wrong variable (`rgw_zonemaster`). Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1925247 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `931b87e830`)	2021-02-10 08:21:41 +01:00
Dimitri Savineau	fa9177d2ce	ceph-mon: add ExecStartPre docker stop to systemd We already do that in the other systemd templates (mgr, mds, etc..) and would present to add workaround in other orchestration tool. This change is for containerized deployment only. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1882724 Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `3749d297c7`)	2021-01-29 11:41:16 -05:00
Guillaume Abrioux	78d9d9df11	rgw: avoid useless call to ceph-rgw since `ceph-rgw` may be called from `ceph-handler` in some contexts we should avoid rerunning it unnecessarily. Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `8617081664`)	2021-01-28 16:37:32 -05:00
Guillaume Abrioux	df98746378	rgw: multisite refact Add the possibility to deploy rgw multisite configuration with a mix of secondary and primary zones on a same rgw node. Before that, on a same node, all instances were either primary zones OR secondary. Now you can define a rgw instance like following: ``` rgw_instances: - instance_name: 'rgw0' rgw_zonemaster: false rgw_zonesecondary: true rgw_zonegroupmaster: false rgw_realm: 'france' rgw_zonegroup: 'zonegroup-france' rgw_zone: paris-00 radosgw_address: "{{ _radosgw_address }}" radosgw_frontend_port: 8080 rgw_zone_user: jacques.chirac rgw_zone_user_display_name: "Jacques Chirac" system_access_key: P9Eb6S8XNyo4dtZZUUMy system_secret_key: qqHCUtfdNnpHq3PZRHW5un9l0bEBM812Uhow0XfB endpoint: http://192.168.101.12:8080 ``` Basically it's now possible to define `rgw_zonemaster`, `rgw_zonesecondary` and `rgw_zonegroupmaster` at the intsance level instead of the whole node level. Also, this commit adds an option `deploy_secondary_zones` (default True) which can be set to `False` in order to explicitly ask the playbook to not deploy secondary zones in case where the corresponding endpoint are not deployed yet. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1915478 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `71a5e666e3`)	2021-01-28 16:37:32 -05:00
Dimitri Savineau	c7d204ce37	ceph-defaults: change default ceph container tag The "latest" ceph container tag references the latest stable release (octopus at the moment). "latest" is an alias on "latest-octopus". On the devel branch we should use "latest-master" tag instead. Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `7d56771975`)	2021-01-22 17:39:39 -05:00
Dimitri Savineau	fdda54eeb4	dashboard: manage password backward compatibility The ceph dashboard changed the way the password are provided via the CLI. This breaks the backward compatibility when using a recent ceph-ansible version with ceph release without that feature. This patch adds tasks for legacy workflow (ceph release without that feature) in both ceph-dashboard role and ceph_dashboard_user module. Signed-off-by: Dimitri Savineau <dsavinea@redhat.com>	2021-01-18 14:46:53 -05:00
Guillaume Abrioux	fb03dfda30	dashboard: configure passwords via stdin Due to recent changes in ceph, the few dashboard passwors must be passed via `-i` Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `ef975ef5ea`)	2021-01-18 14:46:53 -05:00
Guillaume Abrioux	11736265a1	mon: fix cephx disabled deployment Due to missing condition on `cephx` variable, cephx disabled deployments are broken. This commit fixes this. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1910151 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `4af0845702`)	2021-01-18 13:50:57 -05:00
Mike Currin	54f7983be2	Path for ceph config missing in crash template The path where ceph.conf is located (/etc/ceph) missing in the Docker container bind mounts, this throws errors Signed-off-by: Mike Currin <currin@gmail.com> (cherry picked from commit `4cbc9a48c9`)	2021-01-06 16:55:12 +01:00
Guillaume Abrioux	46fac7db28	rgw: support switching from single-site to multisite When collocating rgw with either a mon, mgr or osd, switching from single site to a multisite rgw setup failed because of the handlers triggered between the ansible play of the collocated daemon and the play of the rgw. Since the multisite changes are not yet applied the handlers fail. The idea here is to ensure we run the multisite configuration from the ceph-handler role before the restart happens, this way it won't complain because of non existing multisite configuration. (Note: this is also valid when simply changing a multisite configuration) Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1888630 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `513c8cfe55`)	2021-01-06 10:29:59 -05:00
Fabien Brachere	ba3db6be9f	library: add missing `target_size_ratio` parameter support in ceph_pool module When creating a new pool, target_size_ratio was ignored by ansible module ceph_pool.py. target_size_ratio is now used when pg_autoscale_mode is on. Tests added to library tests. This adds too the use in the role ceph-rgw. Signed-off-by: Fabien Brachere <fabien.brachere@celeste.fr> (cherry picked from commit `4026ba9da1`)	2020-12-16 10:57:33 -05:00
Dimitri Savineau	a2704581b1	ceph-config: fix ceph-volume lvm batch report Since the major ceph-volume lvm batch refactoring, the report value is different. Before the refact, the report was a dict with the OSDs list to be created under the "osds" key. After the refact, the report is a list of dict. Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `827b23353f`)	2020-12-15 17:25:49 -05:00
Dimitri Savineau	d4024eddbb	library: add ceph_osd_flag module This adds ceph_osd_flag ansible module for replacing the command module usage with the ceph osd set/unset commands. Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `5da593604a`)	2020-12-15 14:12:44 +01:00
Dimitri Savineau	1f1ca3ec8a	monitoring: use config_template module for config The alertmanager, grafana and prometheus configuration file are generated with the template module which doesn't allow for using config overrides. Instead we could use the config_template plugin action and add a new variable for overrides (one for each component). With this patch, one should be able to add configuration to prometheus with the following: --- alertmanager_conf_overrides: global: smtp_smarthost: 'localhost:25' ... Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1902999 Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `5a41026347`)	2020-12-14 13:28:12 -05:00
Seena Fallah	e1314de3d9	ceph-osd: use global crush_device_class in lvm_volumes Use global crush_device_class variable if it's not set per OSD Signed-off-by: Seena Fallah <seenafallah@gmail.com> (cherry picked from commit `5e9444fa5c`)	2020-12-14 13:28:00 -05:00
Karl-Heinz Preuß	da7b708636	fix broken ceph-fetch-keys role set fetch_directory variable in default/main.yml instead of using the defaults jinja filter in tasks/main.yml. Fixes: #6072 Signed-off-by: Karl-Heinz Preuß <karl-heinz.preuss@cms.hu-berlin.de> (cherry picked from commit `6ce34ef59f`)	2020-12-14 11:42:50 -05:00
Dimitri Savineau	41f7f9d020	Revert "config: Always use osd_memory_target if set" This reverts commit `4d1fdd2b05`. This breaks the backward compatibility with previous osd_memory_target calculation and we could have a value lower than the minimum value allowed (896M) which causes some ceph commands to fail (like ceph assimilate-conf). Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `aa6e1f20ea`)	2020-12-14 02:41:45 +01:00
Jukka Nousiainen	302fa3b2f8	ceph-mon: No become during gen mon initial keyring Since the backing generate_secret() just hands out urandom output, running as privileged doesn't seem to be required. It's not desireable to provide sudo in some Ansible runner environments. Signed-off-by: Jukka Nousiainen <jukka.nousiainen@csc.fi> (cherry picked from commit `eb7473491b`)	2020-12-07 09:24:37 -05:00
Guillaume Abrioux	6b04f1154f	common: do not use pipefail when not needed Let's discard the ansible lint error 306 and add a "# noqa 306" on tasks where we don't need `set -o pipefail` Fixes: #6090 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `86a8889ee3`)	2020-12-01 20:18:35 -05:00
Guillaume Abrioux	679d3e2d10	osd: add tag on 'wait for all osd to be up' task This allows skipping this task if really desired. Use it carefully. Use it at your own risk. Fixes: #6073 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `5c4ae5356d`)	2020-12-01 11:04:37 +01:00
Guillaume Abrioux	7ab606bac5	iscsigw: remove `--cap-add=all` from `podman run` cmd As of podman `2.0.5`, `--cap-add` and `--privileged` are exclusive options. ``` Nov 30 13:56:30 magna089 podman[171677]: Error: invalid config provided: CapAdd and privileged are mutually exclusive options ``` Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1902149 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `d40dd764e0`)	2020-11-30 16:42:59 -05:00
Guillaume Abrioux	f1ae3dec72	container: remove `--ignore` from `podman rm` command As of podman 2.0.5, `--ignore` param conflicts with `--storage`. ``` Nov 30 13:53:10 magna089 podman[164443]: Error: --storage conflicts with --volumes, --all, --latest, --ignore and --cidfile ``` Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `c68b124ba8`)	2020-11-30 16:42:59 -05:00
Dimitri Savineau	51bea82677	alertmanager/prometheus: fix owner/group Set the owner/group on alertmanager and prometheus directories and files to nobody and nogroup (uid and gid 65534) to avoid permission issues. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1901543 Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `eb452d35bc`)	2020-11-27 14:55:39 -05:00
Guillaume Abrioux	0edbabbf4d	mon: refact initial keyring generation adding monitor is no longer possible because we generate a new mon keyring each time the playbook is run. Fixes: #5864 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `970c6a4ee6`)	2020-11-26 09:12:22 +01:00
Guillaume Abrioux	10551da173	mon: replace `command` task by `copy` We can achieve this task using `copy` module. Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `5ff2ca270f`)	2020-11-26 09:12:22 +01:00
Dimitri Savineau	1cf76da74a	ceph-iscsi: set the pool name in the config file When using a custom pool for iSCSI gateway then we need to set the pool name in the configuration otherwise the default rbd pool name will be used. Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `40a87c4b92`)	2020-11-25 09:19:03 -05:00
Guillaume Abrioux	d86a159a79	osd: ensure /var/lib/ceph/osd/{cluster}-{id} is present This commit ensures that the `/var/lib/ceph/osd/{{ cluster }}-{{ osd_id }}` is present before starting OSDs. This is needed specificly when redeploying an OSD in case of OS upgrade failure. Since ceph data are still present on its devices then the node can be redeployed, however those directories aren't present since they are initially created by ceph-volume. We could recreate them manually but for better user experience we can ask ceph-ansible to recreate them. NOTE: this only works for OSDs that were deployed with ceph-volume. ceph-disk deployed OSDs would have to get those directories recreated manually. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1898486 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `873fc8ec0f`)	2020-11-19 11:52:20 -05:00
Dimitri Savineau	aa302f48de	ceph-facts: fix read osd pool default crush fact We don't need to use run_once on that task when having running monitors otherwise the read task could be skip and the set task will fail. The conditional check 'crush_rule_variable.rc == 0' failed. The error was: error while evaluating conditional (crush_rule_variable.rc == 0): 'dict object' has no attribute 'rc' Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1898856 Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `e150df789e`)	2020-11-18 17:01:05 -05:00
Guillaume Abrioux	acdd43c0e2	containers: modify bindmount option This commit changes the bind mount option for the mount point `/var/lib/ceph` in the systemd template for mon and mgr containers. This is needed in case of collocating mon/mgr with osds using dmcrypt scenario. Once mon/mgr got converted to containers, the dmcrypt layer sub mount is still seen in `/var/lib/ceph`. For some reason it makes the corresponding devices busy so any other container can't open/close it. As a result, it prevents osds from starting properly. Since it only happens on the nodes converted before the OSD play, the idea is to bind mount `/var/lib/ceph` on mon and mgr with the `rshared` option so once the sub mount is unmounted, it is propagated inside the container so it doesn't see that mount point. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1896392 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `f5ba6d9b01`)	2020-11-17 12:27:07 -05:00
Guillaume Abrioux	10dff6888c	container: force rm --storage on ExecStartPre This is a workaround to avoid error like following: ``` Error: error creating container storage: the container name "ceph-mgr-magna022" is already in use by "4a5f674e113f837a0cc561dea5d2cd55d16ca159a647b7794ab06c4c276ef701" ``` that doesn't seem to be 100% reproducible but it shows up after a reboot. The only workaround we came up with at the moment is to run `podman rm --storage <container>` before starting it. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1887716 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `5ba7824c55`)	2020-11-16 16:37:37 -05:00
Benoît Knecht	deaf60316a	ceph-facts: Fix osd_pool_default_crush_rule fact The `osd_pool_default_crush_rule` is set based on `crush_rule_variable`, which is the output of a `grep` command. However, two consecutive tasks can set that variable, and if the second task is skipped, it still overwrites the `crush_rule_variable`, leading the `osd_pool_default_crush_rule` to be set to `ceph_osd_pool_default_crush_rule` instead of the output of the first task. This commit ensures that the fact is set right after the `crush_rule_variable` is assigned, before it can be overwritten. Closes #5912 Signed-off-by: Benoît Knecht <bknecht@protonmail.ch> (cherry picked from commit `c5f7343a2f`)	2020-11-13 10:42:03 -05:00
Gaudenz Steinlin	bb3cfd0481	config: Always use osd_memory_target if set The osd_memory_target variable was only used if it was higher than the calculated value based on the number of OSDs. This is changed to always use the value if it is set in the configuration. This allows this value to be intentionally set lower so that it does not have to be changed when more OSDs are added later. Signed-off-by: Gaudenz Steinlin <gaudenz.steinlin@cloudscale.ch> (cherry picked from commit `4d1fdd2b05`)	2020-11-13 09:27:11 -05:00
Guillaume Abrioux	7b7f20c636	dashboard: change dashboard_grafana_api_no_ssl_verify default value This sets the `dashboard_grafana_api_no_ssl_verify` default value according to the length of `dashboard_crt` and `dashboard_key`. Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `5cadfea42e`)	2020-11-04 11:02:05 -05:00
Guillaume Abrioux	36f550b3b4	dashboard: enable https by default see linked bz for details Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1889426 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `767d3c898e`)	2020-11-04 11:02:05 -05:00
Gaudenz Steinlin	44d43a8e4d	osd: Fix number of OSD calculation If some OSDs are to be created and others already exist the calculation only counted the to be created OSDs. This changes the calculation to take all OSDs into account. Signed-off-by: Gaudenz Steinlin <gaudenz.steinlin@cloudscale.ch> (cherry picked from commit `15044da030`)	2020-11-03 11:36:04 -05:00
Dimitri Savineau	9c70add661	rgw/rbdmirror: use service dump instead of ceph -s The ceph status command returns a lot of information stored in variables and/or facts which could consume resources for nothing. When checking the rgw/rbdmirror services status, we're only using the servicmap structure in the ceph status output. To optimize this, we could use the ceph service dump command which contains the same needed information. This command returns less information and is slightly faster than the ceph status command. $ ceph status -f json \| wc -c 2001 $ ceph service dump -f json \| wc -c 1105 $ time ceph status -f json > /dev/null real 0m0.557s user 0m0.516s sys 0m0.040s $ time ceph service dump -f json > /dev/null real 0m0.454s user 0m0.434s sys 0m0.020s Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `3f9081931f`)	2020-11-03 14:32:09 +01:00
Dimitri Savineau	3bba1fd203	monitor: use quorum_status instead of ceph status The ceph status command returns a lot of information stored in variables and/or facts which could consume resources for nothing. When checking the quorum status, we're only using the quorum_names structure in the ceph status output. To optimize this, we could use the ceph quorum_status command which contains the same needed information. This command returns less information. $ ceph status -f json \| wc -c 2001 $ ceph quorum_status -f json \| wc -c 957 $ time ceph status -f json > /dev/null real 0m0.577s user 0m0.538s sys 0m0.029s $ time ceph quorum_status -f json > /dev/null real 0m0.544s user 0m0.527s sys 0m0.016s Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `88f91d8c12`)	2020-11-03 14:32:09 +01:00
wangxiaotong	b4c1f325a8	osds: use ceph osd stat instead of ceph status Improve the checked way of the OSD created checking process. This replaces the ceph status command by the ceph osd stat command. The osdmap structure isn't needed anymore. $ ceph status -f json \| wc -c 2001 $ ceph osd stat -f json \| wc -c 132 $ time ceph status -f json > /dev/null real 0m0.563s user 0m0.526s sys 0m0.036s $ time ceph osd stat -f json > /dev/null real 0m0.457s user 0m0.411s sys 0m0.045s Signed-off-by: wangxiaotong <wangxiaotong@fiberhome.com> (cherry picked from commit `b9cb0f12e9`)	2020-11-03 14:32:09 +01:00
Guillaume Abrioux	04d47d68fd	common: follow up on #5948 In addition to `f7e2b2c608` Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `371d854a5c`)	2020-11-03 09:43:51 +01:00
Gaudenz Steinlin	2550e44e2f	openstack: use ceph_keyring_permissions by default Otherwise this task fails if no permission is set on the item. Previously the code omited the mode parameter if it was not set, but this was lost with commit `ab370b6ad8`. Signed-off-by: Gaudenz Steinlin <gaudenz.steinlin@cloudscale.ch> (cherry picked from commit `79ff79c422`)	2020-11-02 18:41:53 -05:00
Dimitri Savineau	ec8903f4af	podman: force log driver to journald Since we've changed to podman configuration using the detach mode and systemd type to forking then the container logs aren't present in the journald anymore. The default conmon log driver is using k8s-file. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1890439 Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `16cd183b9c`)	2020-11-02 17:46:39 -05:00
Benoît Knecht	4e15d10e22	ceph-mon: Don't set monitor directory mode recursively After rolling updates performed with `infrastructure-playbooks/rolling_updates.yml`, files located in `/var/lib/ceph/mon/{{ cluster }}-{{ monitor_name }}` had mode 0755 (including the keyring), making them world-readable. This commit separates the task that configured permissions recursively on `/var/lib/ceph/mon/{{ cluster }}-{{ monitor_name }}` into two separate tasks: 1. Set the ownership and mode of the directory itself; 2. Recursively set ownership in the directory, but don't modify the mode. Signed-off-by: Benoît Knecht <bknecht@protonmail.ch> (cherry picked from commit `0d76826bbb`)	2020-11-02 17:03:04 -05:00
Dimitri Savineau	fa83929b8e	ceph-handler: fix curl ipv6 command with rgw When using the curl command with ipv6 address and brackets then we need to use the -g option otherwise the command fails. $ curl http://[fdc2:328:750b:6983::6]:8080 curl: (3) [globbing] error: bad range specification after pos 9 Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `cdb7b09cd7`)	2020-11-02 16:34:51 -05:00

1 2 3 4 5 ...

2857 Commits (1b49988e4aa311559287eb51dc4040584ad1472a)