ceph-ansible

Commit Graph

Author	SHA1	Message	Date
Dimitri Savineau	c39e7cb151	alertmanager: allow disable dashboard tls verify When using self-signed/untrusted CA certificates, alertmanager displays an error in logs. With this commit this should make those messages disappear. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1936299 Co-authored-by: Guillaume Abrioux <gabrioux@redhat.com> Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `9f77b929d1`)	2021-07-26 13:19:13 -04:00
Guillaume Abrioux	72bbc8285e	dashboard: support dedicated network for the dashboard This introduces a new variable `dashboard_network` in order to support deploying the dashboard on a different subnet. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1927574 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `f4f73b6197`)	2021-07-26 13:19:03 -04:00
Dimitri Savineau	00e0ebc911	multisite: use node fqdn for endpoints when https When the rgw_multisite_proto variable is set to https then we shoudn't use the IP address in the zone endpoints list but the node FQDN to match the TLS certificate CN. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1965504 Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `ad05a08160`)	2021-07-26 17:54:13 +02:00
Dimitri Savineau	eba580320c	ceph-mgr: don't install dashboard pkg by default This is a partial backport of `2547ab60`. We are currently installing the ceph-mgr-dashboard package even if the dashboard_enabled variable is set to false. Signed-off-by: Dimitri Savineau <dsavinea@redhat.com>	2021-07-26 17:50:42 +02:00
Dimitri Savineau	8d58c50f45	ceph-mgr: move mgr module list to common Populating the ceph_mgr_modules list in the mgr_modules doesn't make sense since that file is only executed if the list isn't empty or we're using the dashboard. Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `cd06e7c046`)	2021-07-26 17:50:35 +02:00
Dimitri Savineau	364186a86e	ceph-nfs: allow overriding NFS_CORE_PARAM We already have config override variables for existing block (like ganesha_ceph_export_overrides, ganesha_log_overrides, etc...) or a global one (ganesha_conf_overrides) but redefining the NFS_CORE_PARAM block in that variable will erase all previous values (currently only Bind_Addr). ganesha_core_param_overrides: \| Enable_UDP = false; NFS_Port = 2050; Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1941775 Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `9817d29543`)	2021-07-26 17:50:05 +02:00
Dimitri Savineau	ddc3df9f9a	ceph-facts: move device facts to its own file Instead of reusing the condition 'inventory_hostname in groups[osds]' on each device facts tasks then we can move all the tasks into a dedicated file and set the condition on the import_tasks statement. Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `d704b05e52`)	2021-07-26 17:49:03 +02:00
Dimitri Savineau	50447e89fb	ceph-validate: check logical volumes We currently don't check if the logical volume used in lvm_volumes list for either bluestore data/db/wal or filestore data/journal exist. We're only doing this on raw devices for batch scenario. Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `55bca07cb6`)	2021-07-26 17:49:03 +02:00
Dimitri Savineau	ceca225344	ceph-validate: check db/journal/wal devices too When using dedicated devices for db/journal/wal objecstore with ceph-volume lvm batch then we should also validate that those devices exist and don't use a gpt partition table in addition of the devices and lvm_volume.data variables. Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `808e7106de`)	2021-07-26 17:49:03 +02:00
Dimitri Savineau	fe070fc19d	ceph-validate: use root device from ansible_mounts Instead of using findmnt command to find the device associated to the root mount point then we can use the ansible_mounts fact. Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `7e50380f7f`)	2021-07-26 17:49:03 +02:00
Dimitri Savineau	f317df92ac	ceph-validate: do not resolve devices This is already done in the ceph-facts role. Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `0df99dda8d`)	2021-07-26 17:49:03 +02:00
Dimitri Savineau	c67bfe84eb	ceph-validate: check block presence first Instead of doing two parted calls we can check first if the device exist and then test the partition table. Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `14d458b3b4`)	2021-07-26 17:49:03 +02:00
Dimitri Savineau	5ef1d630d8	ceph-validate: check devices from lvm_volumes `2888c08` introduced a regression as the check_devices tasks file was only included based on the devices variable. But that file also validate some devices from the lvm_volumes variable. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1906022 Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `ac0342b72e`)	2021-07-26 17:49:03 +02:00
Dimitri Savineau	4695df6d2b	monitoring: use config_template module for config The alertmanager, grafana and prometheus configuration file are generated with the template module which doesn't allow for using config overrides. Instead we could use the config_template plugin action and add a new variable for overrides (one for each component). With this patch, one should be able to add configuration to prometheus with the following: --- alertmanager_conf_overrides: global: smtp_smarthost: 'localhost:25' ... Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1902999 Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `5a41026347`)	2021-07-26 17:47:51 +02:00
Dimitri Savineau	17b9ff03d2	common: fix py2 pool_list from_json when skipped When using python 2 and the task with a loop is skipped then it generates an error. Unexpected templating type error occurred on ({{ (pool_list.stdout \| from_json)['pools'] }}): expected string or buffer Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `cf6e33346e`)	2021-07-21 09:54:46 -04:00
Guillaume Abrioux	f7882bbc02	common: disable/enable pg_autoscaler The PG autoscaler can disrupt the PG checks so the idea here is to disable it and re-enable it back after the restart is done. Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `13036115e2`)	2021-07-21 09:40:18 -04:00
Neelaksh Singh	5213612eaf	Sensitive key data now hidden in output log Fixes: #6529 Signed-off-by: Neelaksh Singh <neelaksh48@gmail.com> (cherry picked from commit `d18a9860cd`)	2021-07-12 09:43:12 +02:00
Dimitri Savineau	58dddf586e	Revert "ceph-validate: check devices from lvm_volumes" This reverts commit `3557497336`. Signed-off-by: Dimitri Savineau <dsavinea@redhat.com>	2021-07-07 17:19:35 +02:00
Dimitri Savineau	a684a26428	Revert "ceph-validate: check block presence first" This reverts commit `4f89cdcd45`. Signed-off-by: Dimitri Savineau <dsavinea@redhat.com>	2021-07-07 17:19:35 +02:00
Dimitri Savineau	57f9553798	Revert "ceph-validate: do not resolve devices" This reverts commit `2020b1310c`. Signed-off-by: Dimitri Savineau <dsavinea@redhat.com>	2021-07-07 17:19:35 +02:00
Dimitri Savineau	bc570619b6	Revert "ceph-validate: use root device from ansible_mounts" This reverts commit `b1542fd340`. Signed-off-by: Dimitri Savineau <dsavinea@redhat.com>	2021-07-07 17:19:35 +02:00
Dimitri Savineau	e9123dda35	Revert "ceph-validate: check db/journal/wal devices too" This reverts commit `d6f3e6eac3`. Signed-off-by: Dimitri Savineau <dsavinea@redhat.com>	2021-07-07 17:19:35 +02:00
Dimitri Savineau	c096ec4033	Revert "ceph-validate: check logical volumes" This reverts commit `d7cefe0536`. Signed-off-by: Dimitri Savineau <dsavinea@redhat.com>	2021-07-07 17:19:35 +02:00
Dimitri Savineau	b82f4edb38	Revert "ceph-facts: move device facts to its own file" This reverts commit `9f1ec38bbf`. Signed-off-by: Dimitri Savineau <dsavinea@redhat.com>	2021-07-07 17:19:35 +02:00
Guillaume Abrioux	928d7c75a4	dashboard: remove "certificate is valid for" error When deploying dashboard with ssl certificates generated by ceph-ansible, we enforce the CN to 'ceph-dashboard' which can makes application such alertmanager complain like following: `err="Post https://mgr0:8443/api/prometheus_receiver: x509: certificate is valid for ceph-dashboard, not mgr0" context_err="context deadline exceeded"` The idea here is to add alternative names matching all mgr/mon instances in the certificate so this error won't appear in logs. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1978869 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `72a0336c71`)	2021-07-07 17:19:22 +02:00
Dimitri Savineau	9f1ec38bbf	ceph-facts: move device facts to its own file Instead of reusing the condition 'inventory_hostname in groups[osds]' on each device facts tasks then we can move all the tasks into a dedicated file and set the condition on the import_tasks statement. Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `d704b05e52`)	2021-07-05 18:03:43 +02:00
Dimitri Savineau	d7cefe0536	ceph-validate: check logical volumes We currently don't check if the logical volume used in lvm_volumes list for either bluestore data/db/wal or filestore data/journal exist. We're only doing this on raw devices for batch scenario. Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `55bca07cb6`)	2021-07-05 18:03:43 +02:00
Dimitri Savineau	d6f3e6eac3	ceph-validate: check db/journal/wal devices too When using dedicated devices for db/journal/wal objecstore with ceph-volume lvm batch then we should also validate that those devices exist and don't use a gpt partition table in addition of the devices and lvm_volume.data variables. Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `808e7106de`)	2021-07-05 18:03:43 +02:00
Dimitri Savineau	b1542fd340	ceph-validate: use root device from ansible_mounts Instead of using findmnt command to find the device associated to the root mount point then we can use the ansible_mounts fact. Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `7e50380f7f`)	2021-07-05 18:03:43 +02:00
Dimitri Savineau	2020b1310c	ceph-validate: do not resolve devices This is already done in the ceph-facts role. Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `0df99dda8d`)	2021-07-05 18:03:43 +02:00
Dimitri Savineau	4f89cdcd45	ceph-validate: check block presence first Instead of doing two parted calls we can check first if the device exist and then test the partition table. Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `14d458b3b4`)	2021-07-05 18:03:43 +02:00
Dimitri Savineau	3557497336	ceph-validate: check devices from lvm_volumes `2888c08` introduced a regression as the check_devices tasks file was only included based on the devices variable. But that file also validate some devices from the lvm_volumes variable. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1906022 Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `ac0342b72e`)	2021-07-05 18:03:43 +02:00
Dimitri Savineau	04c18710ac	prometheus: fix prometheus target url The prometheus service isn't binding on localhost. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1933560 Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `1d56818658`)	2021-07-02 14:37:56 -04:00
Guillaume Abrioux	ff2043f92c	ceph_key: handle error in a better way When calling the `ceph_key` module with `state: info`, if the ceph command called fails, the actual error is hidden by the module which makes it pretty difficult to troubleshoot. The current code always states that if rc is not equal to 0 the keyring doesn't exist. `state: info` should always return the actual rc, stdout and stderr. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1964889 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `d58500ade0`)	2021-07-02 14:01:52 +02:00
Dimitri Savineau	77f32a3302	container: set tcmalloc value by default All ceph daemons need to have the TCMALLOC_MAX_TOTAL_THREAD_CACHE_BYTES environment variable set to 128MB by default in container setup. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1970913 Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `9758e3c513`)	2021-07-01 15:46:19 +02:00
Boris Ranto	a6cf646e45	dashboard: Add new prometheus alert It was requested for us to update our alerting definitions to include a slow OSD Ops health check. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1951664 Signed-off-by: Boris Ranto <branto@redhat.com> (cherry picked from commit `2491d4e004`)	2021-07-01 09:37:37 +02:00
Guillaume Abrioux	0fa1cf0cdf	nfs: do no copy client.bootstrap-rgw when using mds There's no need to copy this keyring when using nfs with mds Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `8dbee99882`)	2021-06-17 08:15:53 +02:00
VasishtaShastry	7bc9e391cb	Container: Fixing service name lvm2-lvmetad Playbook failing saying: msg: 'Could not find the requested service lvmetad: host' Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1955040 Signed-off-by: VasishtaShastry <vipin.indiasmg@gmail.com> (cherry picked from commit `e49c38f8b7`)	2021-06-17 08:14:54 +02:00
Guillaume Abrioux	98eb93db3e	multisite: fix bug during switch2containers When running the switch-to-containers playbook with multisite enabled, the fact "rgw_instances" is only set for the node being processed (serial: 1), the consequence of that is that the set_fact of 'rgw_instances_all' can't iterate over all rgw node in order to look up each 'rgw_instances_host'. Adding a condition checking whether hostvars[item]["rgw_instances_host"] is defined fixes this issue. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1967926 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `8279d14d32`)	2021-06-17 07:20:37 +02:00
Guillaume Abrioux	611494b88f	rolling_update: fix mon+rgw/multisite collocation When monitors and rgw are collocated with multisite enabled, the rolling_update playbook fails because during the workflow, we run some radosgw-admin commands very early on the first mon even though this is the monitor being upgraded, it means the container doesn't exist since it was stopped. This block is relevant only for scaling out rgw daemons or initial deployment. In rolling_update workflow, it is not needed so let's skip it. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1970232 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `f7166cccbf`)	2021-06-14 13:59:16 +02:00
Guillaume Abrioux	71764c3440	tests: disable test_mgr_dashboard_is_listening Due to a recent commit that has introduced a regression in ceph, this test is failing. Temporarily disabling it to unblock the CI. Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `2e19d1705e`)	2021-06-07 15:12:43 +02:00
Guillaume Abrioux	d00e4b3d9e	dashboard: set cookie_secure in grafana When using grafana behind https `cookie_secure` should be set to `true`. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1966880 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `4daed1f137`)	2021-06-07 15:12:43 +02:00
Guillaume Abrioux	f3b992328a	dashboard: fix rgw user creation When deploying dashboard in a cluster with rgw multisite deployed. Due to the last rgw multisite refactor, we now expect the variable `rgw_zonemaster` to be defined in the dict `rgw_instances`. The idea here is to create that user on the cluster as soon as we have 1 `rgw_zonemaster` set to `true` in `rgw_instances`. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1964995 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2021-06-03 17:07:02 +02:00
Guillaume Abrioux	d2bd0961d4	crash: fix --limit deployments (containers) ceph-crash deployments is broken when ceph-ansible playbook is called with --limit in containerized contexts since we don't set `container_exec_cmd` on the first monitor. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1964835 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `10ed26f14d`)	2021-05-26 18:51:48 +02:00
Guillaume Abrioux	a391dad8e1	dashboard: fix typo introduced during backport during backport of `c8b92deba1` the pattern should have been s/monitoring_group_name/grafana_server_group_name/ Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1964907 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `ac0a5c1e68`)	2021-05-26 18:51:18 +02:00
Guillaume Abrioux	8b4eb0f108	prometheus: enforce osd nodes in templates When osd nodes are collocated in the clients group (HCI context for instance), the current logic will exclude osd nodes since they are present in the client group. The best fix would be to exclude clients node only when they are not member of another group but for now, as a workaround, we can enforce the addition of osd nodes to fix this specific case. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1947695 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `664dae0564`)	2021-05-26 10:57:51 +02:00
Guillaume Abrioux	a27761855b	container: conditionnally disable lvmetad Enabling lvmetad in containerized deployments on el7 based OS might cause issues. This commit make it possible to disable this service if needed. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1955040 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2021-05-25 16:51:04 +02:00
Brad Hubbard	5d3a46e6fd	Make sure the repo url contains the correct arch We can end up with an arm only repo unless we are specific about the architecture we require. Brings the deb code in line with the rpm equivalent. Signed-off-by: Brad Hubbard <bhubbard@redhat.com> (cherry picked from commit `267cce9e83`)	2021-05-17 08:57:19 +02:00
Guillaume Abrioux	6999118fb6	validate: check virtual_ips variable This commit checks the length of `virtual_ips` doesn't exceed the length of `groups[rgwloadbalancer_group_name]`. It also ensure this variable is defined when `groups[rgwloadbalancer_group_name]` contains at least one node. Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `3b63e0649c`)	2021-05-05 09:56:42 +02:00
Benoît Knecht	97066a1ebc	ceph-rgw-loadbalancer: Fix keepalived master selection While `2ca33641` fixed a bug in the way the `keepalived.conf.j2` template matched hostnames to set the VRRP `MASTER`/`BACKUP` states, it also introduced a regression in the case where `virtual_ips` is a list of more than one IP address. The previous behavior would result in each host in the `rgwloadbalancers` group to be `MASTER` for one of the `virtual_ips`, but the new behavior caused the first host to be `MASTER` for all the IP address in `virtual_ips`. This commit restores the original behavior. Signed-off-by: Benoît Knecht <bknecht@protonmail.ch> (cherry picked from commit `2bede4762e`)	2021-05-05 09:56:42 +02:00

1 2 3 4 5 ...

2762 Commits (c39e7cb1516eebd57a2d5ca99e8e5aeefc77a980)