ceph-ansible

Commit Graph

Author	SHA1	Message	Date
Guillaume Abrioux	a27761855b	container: conditionnally disable lvmetad Enabling lvmetad in containerized deployments on el7 based OS might cause issues. This commit make it possible to disable this service if needed. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1955040 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2021-05-25 16:51:04 +02:00
Brad Hubbard	5d3a46e6fd	Make sure the repo url contains the correct arch We can end up with an arm only repo unless we are specific about the architecture we require. Brings the deb code in line with the rpm equivalent. Signed-off-by: Brad Hubbard <bhubbard@redhat.com> (cherry picked from commit `267cce9e83`)	2021-05-17 08:57:19 +02:00
Guillaume Abrioux	6999118fb6	validate: check virtual_ips variable This commit checks the length of `virtual_ips` doesn't exceed the length of `groups[rgwloadbalancer_group_name]`. It also ensure this variable is defined when `groups[rgwloadbalancer_group_name]` contains at least one node. Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `3b63e0649c`)	2021-05-05 09:56:42 +02:00
Benoît Knecht	97066a1ebc	ceph-rgw-loadbalancer: Fix keepalived master selection While `2ca33641` fixed a bug in the way the `keepalived.conf.j2` template matched hostnames to set the VRRP `MASTER`/`BACKUP` states, it also introduced a regression in the case where `virtual_ips` is a list of more than one IP address. The previous behavior would result in each host in the `rgwloadbalancers` group to be `MASTER` for one of the `virtual_ips`, but the new behavior caused the first host to be `MASTER` for all the IP address in `virtual_ips`. This commit restores the original behavior. Signed-off-by: Benoît Knecht <bknecht@protonmail.ch> (cherry picked from commit `2bede4762e`)	2021-05-05 09:56:42 +02:00
Guillaume Abrioux	2d59f4579b	update: fix ceph-crash stop task This is a workaround for an issue in ansible. When trying to stop/mask/disable this service in one task, the stop didn't actually happen, the task doesn't fail but for some reason the container is still present and running. Then the task starting the service in the role ceph-crash fails because it can't start the container since it's already running with the same name. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1955393 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `3db1ea7ec4`)	2021-05-05 09:47:32 +02:00
Seena Fallah	4f22dc477e	ceph-osd: allow to use ceph_tcmalloc_max_total_thread_cache for bluestore TCMALLOC_MAX_TOTAL_THREAD_CACHE_BYTES is for both bluestore and filestore Signed-off-by: Seena Fallah <seenafallah@gmail.com> (cherry picked from commit `41295f0ef6`)	2021-04-29 07:34:55 +02:00
Benoît Knecht	104ba407f2	ceph-mon: Fix check mode for deploy monitor tasks Skip the `get initial keyring when it already exists` task when both commands whose `stdout` output it requires have been skipped (e.g. when running in check mode). Signed-off-by: Benoît Knecht <bknecht@protonmail.ch> (cherry picked from commit `e98d9b70bd`)	2021-04-28 10:00:30 +02:00
Guillaume Abrioux	650964a8c7	fs2bs: add a final play This removes the fact `skipped_nodes` which is useless when we run with `--limit` since it gets reset when a new iteration is made. Instead, let's print within a final play which node has been skipped reusing the `skip_this_node` fact. Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `3d4267051f`)	2021-04-28 08:55:34 +02:00
Francesco Pantano	86091f5ba3	Config the monitoring stack components api urls using a VIP When dashboard_frontend_vip is provided, all the services should be configured using the related VIP. A new VIP variable is added for both prometheus and alertmanager: we're already able to properly config the grafana vip using dashboard_frontend_vip variable. This change adds the same variable for both prometheus and alertmanager. Signed-off-by: Francesco Pantano <fpantano@redhat.com> (cherry picked from commit `441651638d`)	2021-04-28 08:54:21 +02:00
Guillaume Abrioux	871ad6d71a	osd: always allow setting target_size_ratio We shouldn't prevent from setting target_size_ratio when the autoscaler is set to 'warn'. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1906305 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2021-04-15 14:47:38 +02:00
Benoît Knecht	9a286e273a	ceph-rgw-loadbalancer: Fix rgw_ports fact The `set_fact rgw_ports` task was failing due to a templating error, because `hostvars[item].rgw_instances` is a list, but it was treated as if it was a dictionary. Another issue was the fact that the `unique` filter only applied to the list being appended to `rgw_ports` instead of the entire list, which means it was possible to have duplicate items. Lastly, `rgw_ports` would have been a list of integers, but the `seport` module expects a list of strings. This commit fixes all of the issues above, allowing the `ceph-rgw-loadbalancer` role to work on systems with SELinux enabled. Signed-off-by: Benoît Knecht <bknecht@protonmail.ch> (cherry picked from commit `c078513475`)	2021-04-15 13:21:07 +02:00
Guillaume Abrioux	b4340b71d9	switch-to-containers: only chown corresponding files When collocating daemons, if we chown all files under `/var/lib/ceph` it can cause issues for the collocated daemons that wouldn't have been migrated yet. This commit makes the playbook chown only the files corresponding to the daemon being migrated. Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `ddbc11c4a9`)	2021-04-15 05:24:44 +02:00
Guillaume Abrioux	ddd7c42c2b	container/systemd: ensure /var/log/ceph exists This adds a `ExecStartPre=-/usr/bin/mkdir -p /var/log/ceph` in all systemd service templates for all ceph daemon. This is specific to RHCS after a Leapp upgrade is done. Indeed, the `/var/log/ceph` seems to be removed after the upgrade. In order to work around this issue let's ensure the directory is present before trying to start the containers with podman. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1949489 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `bab403b603`)	2021-04-14 20:46:09 +02:00
Guillaume Abrioux	4db92dae59	rbdmirror: add retries/until when configuring mirroring `configure_mirroring.yml` is called right after the daemon is started. Sometimes, it can happen the first task in `configure_mirroring.yml` is run while the daemon isn't yet ready, adding a retries/until on that task should help to avoid causing the playbook to fail. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1944996 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `b1e7e1ad0f`)	2021-04-14 16:13:01 +02:00
Guillaume Abrioux	3ef9690cd1	docker2podman: skip some role imports from handler when running docker-to-podman playbook, there's no need to call `ceph-config` and `ceph-rgw` from the role `ceph-handler`. It can even have side effects when coming from a baremetal cluster that was previously migrated using the switch-to-containers playbook. Indeed it might complain about missing .target systemd unit since they are removed during that migration. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1944999 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `70f19be367`)	2021-04-12 13:30:31 +02:00
Guillaume Abrioux	c0c90c6747	docker2podman: add documentation/header this adds a small documentation in the header of the playbook in order to explain what is the goal of this playbook. Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `36b4227dcd`)	2021-04-12 09:45:20 +02:00
Guillaume Abrioux	3bf2c45123	switch_to_containers: support iscsigws migration This adds the iscsigws migration to containers. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=<bz-number> Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `2c74c27321`)	2021-04-09 15:28:27 +02:00
Guillaume Abrioux	f47da73a8a	common: selinux tasks related refactor This moves some task from the `ceph-nfs` role in `ceph-common` since some of them are needed in `ceph-rgwloadbalancer` role. This avoids duplicated tasks. Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `d0442d81b9`)	2021-04-06 15:09:00 +02:00
Guillaume Abrioux	3bfa0772e2	rgw-loadbalancers: add all rgw_ports to http_port_t type This adds all rgw ports to the http_port_t selinux type so it allows haproxy to connect to those ports in order to avoid AVC. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1923890 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `6bbb90198b`)	2021-04-06 15:09:00 +02:00
kalebskeithley	3ea39b5db3	rgw-loadbalancer: Update haproxy.cfg.j2 haproxy gets an AVC when configured to connect to port 8081 This commit adds a snippet regarding haproxy in a selinux environment Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1923890 Signed-off-by: Kaleb S KEITHLEY <kkeithle@redhat.com> (cherry picked from commit `9e7f22a071`)	2021-04-06 15:09:00 +02:00
Guillaume Abrioux	b79c2070f4	nfs: set idmap config for Ceph-NFS Currently NFS Ganesha (ceph-nfs) consumes /etc/idmapd.conf, which controls mapping of user/owner identities under NFSv4+. With containerized service deployment, this file is an immutable part of the container image and cannot be modified. Here we provide group variables, and a taskk and templates for the ceph-nfs role, to set the path of the idmap configuration file and to make the most common adjustment to the contents of that file -- namely to set the 'Domain'. We default the path to /etc/ganesha/idmap.conf so that we will not conflict with /etc/idmapd.conf on the controller nodes where ganesha runs. NFSv4 clients, as used for example by the Cinder NFS driver, consume /etc/idmapd.conf and may require different settings than what is wanted for NFS Ganesha. Additionally, because we already bind /etc/ganesha from the host into the ceph-nfs container, the file NFS Ganesha consumes will no longer be an immutable part of the container. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1925646 Signed-off-by: Tom Barron tpb@dyncloud.net Co-Authored-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `2db2208e40`)	2021-04-02 13:18:52 +02:00
Guillaume Abrioux	b2cf677b71	dashboard: support prometheus storage.tsdb.retention.time parameter This commit adds the parameter `--storage.tsdb.retention.time` to the prometheus systemd unit template. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1928000 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `b60c61ce45`)	2021-04-02 13:17:59 +02:00
Guillaume Abrioux	5fd299e358	update: followup on `07029e1` Playbook must fail anyway, the `rescue` block has been introduced for unmasking the unit after the playbook has failed. Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `e9ddb972fe`)	2021-03-29 15:22:23 +02:00
Guillaume Abrioux	82b934cfc1	rolling_update: unmask monitor service after a failure if for some reason the playbook fails after the service was stopped, disabled and masked and before it got restarted, enabled and unmasked, the playbook leaves the service masked and which can make users confused and forces them to unmask the unit manually. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1917680 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `07029e1bf1`)	2021-03-29 15:22:23 +02:00
Guillaume Abrioux	653d180ec0	defaults: add a comment about `igw_network` This add a quick documentation in ceph-defaults about `igw_network` Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `c5728bdc63`)	2021-03-29 11:24:28 +02:00
Guillaume Abrioux	fe47a02134	dashboard: support igw nodes with dedicated subnet This adds the possibility to deploy the dashboard with igw nodes using a dedicated subnet. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1926170 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `c33de174f1`)	2021-03-26 21:26:14 +01:00
VasishtaShastry	58a28656ff	Peer addition won't be skipped if remote is not in peer rbd-mirroring is not configured as adding peer is getting skipped. Peer addition should not get skipped if its not added already Closes - https://bugzilla.redhat.com/show_bug.cgi?id=1942444 Signed-off-by: VasishtaShastry <vipin.indiasmg@gmail.com> (cherry picked from commit `006998e804`)	2021-03-26 19:14:35 +01:00
Guillaume Abrioux	a8420d41c6	update: stop ceph-crash service before upgrading This adds the missing service stop task for ceph-crash upgrade workflow. It should have been added through commit `15872e3db1e342238636bc9c8e1aef6bd1d3dcd8` in stable-4.0 but at the time we backported this patch ceph-crash wasn't implemented yet so the ceph-crash related content in this patch was removed. Then, ceph-crash has been implemented later so we are still missing this part of the patch in stable-4.0. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1943471 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2021-03-26 16:18:50 +01:00
Guillaume Abrioux	a65968a9d1	tests: pin ruamel.yaml version 0.17.0 which was released today (03/26/2021) breaks ansible-lint execution with py2.7. From https://pypi.org/project/ruamel.yaml we can read: > The 0.16.13 release was the last that will tested to be working on Python 2.7. Let's enforce the version on 0.16.13 when running with py2.7 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2021-03-26 14:42:27 +01:00
Ali Maredia	ba6fa5c959	docs: rgw multisite docs with new rgw_instances config Docs reflect that each instance of `rgw_instances` can now take rgw_zonemaster, rgw_zonesecondary, rgw_zonegroupmaster, rgw_multisite_proto. Signed-off-by: Ali Maredia <amaredia@redhat.com> (cherry picked from commit `a59bc2da3b`)	2021-03-26 07:43:02 +01:00
Guillaume Abrioux	9780490b2f	convert some missed `ansible_`` calls to `ansible_facts['']` This converts some missed calls to `ansible_*` that were missed in initial PR #6312 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `0163ecc924`)	2021-03-26 00:16:58 +01:00
Alex Schultz	6229b3bdba	Disable facts by default in ansible.cfg As a continuation of `a7f2fa73e6`, this change switches fact injection to off by default in the provided ansible.cfg. Signed-off-by: Alex Schultz <aschultz@redhat.com> (cherry picked from commit `db031a4993`)	2021-03-26 00:16:58 +01:00
Alex Schultz	7ddbe74712	Use ansible_facts It has come to our attention that using ansible_* vars that are populated with INJECT_FACTS_AS_VARS=True is not very performant. In order to be able to support setting that to off, we need to update the references to use ansible_facts[<thing>] instead of ansible_<thing>. Related: ansible#73654 Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1935406 Signed-off-by: Alex Schultz <aschultz@redhat.com> (cherry picked from commit `a7f2fa73e6`)	2021-03-26 00:16:58 +01:00
Guillaume Abrioux	697e5823f3	library: drop ceph_facts This is never called in the playbook and seems unmaintained. Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `b01f16e835`)	2021-03-26 00:07:43 +01:00
Ken Dreyer	4aadd659e2	README-MULTISITE: fix typos This commit fixes some typos in MULTISITE documentation. Signed-off-by: Ken Dreyer <ktdreyer@redhat.com> (cherry picked from commit `63a246db41`)	2021-03-26 00:07:04 +01:00
Guillaume Abrioux	93defc4f4b	tests: switch to quay.ceph.io for dashboard images for some reason, `quay.io/app-sre/grafana` no longer exist. as a workaround, all dashboard related images have been mirrored on quay.ceph.io. Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `c90b0985e5`)	2021-03-25 14:11:11 +01:00
Guillaume Abrioux	7fd332e7fe	iscsi: fetch right repo from shaman due to recent changes in shaman, we must fetch the right repo by filtering on the desired architecture. Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `5801171b37`)	2021-03-25 14:11:11 +01:00
Guillaume Abrioux	bd6cc79fa0	tests: fix `test_rgw_is_up` test The data structure seems to have been modified in ceph@master (quincy). This commit update the test accordingly. Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `b8080bac41`)	2021-03-25 14:11:11 +01:00
Guillaume Abrioux	358ea3853a	tests: fix `test_nfs_is_up` test the data structure seems to have been modified in ceph@master (quincy). This commit update the test accordingly. Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `7e1db0b599`)	2021-03-25 14:11:11 +01:00
Guillaume Abrioux	b46d2bf0a6	ceph_volume: fix bug in `is_lv()` This function makes the `ceph_volume` module be not idempotent in containerized context because it tries to run a container and bindmount directories that no longer exist. In that case, the `lvs` command being executed returns something different than `0` so we can't call `json.loads(out)['report'][0]['lv']` since it might throw an python error. The idea is to return `True` only if `rc` is equal to `0` and `len(result)` is greater than `0`, which means the command matched an LV. Fixes: #6284 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `ed79bc7a4e`)	2021-03-25 14:11:11 +01:00
Guillaume Abrioux	2cd8c3637c	fix 'command -v' tasks `command -v` is a bash script which needs a shell to run. Fixes: #6325 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `14c472707c`)	2021-03-22 13:53:11 +01:00
Guillaume Abrioux	bbf8b2fdf6	facts: fix nfs/external cluster scenario These tasks shouldn't be run when at least 1 monitor isn't present in the inventory. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1937997 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `ccd1cbb732`)	2021-03-18 06:41:00 +01:00
Guillaume Abrioux	dc2a11ce3f	config: reset num_osds When collocating OSDs with other daemon, `num_osds` is incorrectly calculated because `ceph-config` is called multiple times. Indeed, the following code: ``` num_osds: "{{ lvm_list.stdout \| default('{}') \| from_json \| length \| int + num_osds \| default(0) \| int }}" ``` makes `num_osds` be incremented each time `ceph-config` is called. We have to reset it in order to get the correct number of expected OSDs. Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `31a0f2653d`)	2021-03-17 17:35:52 +01:00
Guillaume Abrioux	8b86b2ede3	tests: increase nb of rerun in pytest In order to avoid false positive in the CI that I've been unable to reproduce. Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `f7fd1c2298`)	2021-03-12 17:52:00 +01:00
Matthew Vernon	ce25fc74eb	Docs: fix some typos While working on the previous PR, I found a couple of typos in the docs. This fixes those. Signed-off-by: Matthew Vernon <mv3@sanger.ac.uk> (cherry picked from commit `8b1474ab75`)	2021-03-12 09:36:11 +01:00
Dimitri Savineau	6921aafb2b	debian/uca: remove the handler notification The "update apt cache" in the ceph-handler role was never called and the handler trigger after adding the uca repository doesn't exist at all. Instead of using a handler for that we can just set the update_cache parameter to true like the other apt_repository tasks. Resolve merge conflict from cherry-picking this commit. Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `09d6706697`)	2021-03-11 22:06:11 +01:00
Guillaume Abrioux	e6447bdc2b	library: do not always add --yes in batch mode When asking `ceph-volume` to report only in `lvm batch` context, there's a bug described in bz1896803 [1] when `--yes` is passed (which by the way isn't necessary with `--report`). This commit ensure `--yes` isn't passed to `ceph-volume` when `--report` is used. [1] https://bugzilla.redhat.com/show_bug.cgi?id=1896803 Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1896803 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `fe6d6ba622`)	2021-03-11 13:53:06 +01:00
Guillaume Abrioux	0d0723298f	purge: rm service-cid files This commit makes sure purge playbooks remove those file if for any reason they have been left. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1920900 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `b9dd253a4f`)	2021-03-11 13:52:48 +01:00
Guillaume Abrioux	932abbc8cf	switch2container: do not serialize the ceph-crash migration There's no need to slow down the playbook execution time by migrating all the `ceph-crash` instances in a serial way. Let's remove the `serial: 1` so the migration is achieved in a parallel way. Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `980a5a7df4`)	2021-03-11 13:52:39 +01:00
Dimitri Savineau	8f26ffdbac	rolling_update: enforce ceph-container-engine When running the rolling_update.yml playbook and adding the dashboard component in the same time then the requirement (like container packages) aren't installed. This could lead to a failure in case of using authentication on the container registry because the playbook will try to login on the registry but podman/docker aren't yet installed. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1903504 Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1918650 Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `48a456dc8c`)	2021-03-11 13:52:21 +01:00

1 2 3 4 5 ...

5452 Commits (91834f6c9f86214a87a71cae1e3dd9d267b0d7a7) All Branches Search

5452 Commits (91834f6c9f86214a87a71cae1e3dd9d267b0d7a7)

All Branches