ceph-ansible

Commit Graph

Author	SHA1	Message	Date
Guillaume Abrioux	041435e1e3	rbd-mirror: follow up on recent rbd-mirror refactor - ensure /var/lib/ceph/bootstrap-rbd-mirror exists - always install ceph-base on rbdmirror nodes (otherwise, ceph-crash isn't present) Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2022-08-02 10:35:33 +02:00
Teoman ONAY	0c50bfac98	Set ceph_rbd_mirror_pool default value Signed-off-by: Teoman ONAY <tonay@redhat.com>	2022-08-02 10:35:33 +02:00
Teoman ONAY	cef1636f70	Playbook fails when using --limit to install new MDS "set_fact container_run_cmd" is not set when using --limit on MDS as facts were not run on first MON. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=2111017 Signed-off-by: Teoman ONAY <tonay@redhat.com>	2022-08-02 10:35:33 +02:00
Guillaume Abrioux	b74ff6e22c	rbd-mirror: major refactor - Use config-key store to add cluster peer. - Support multiple pools mirroring. Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2022-07-29 17:33:25 +02:00
Guillaume Abrioux	cf4a430d0b	config: followup on `8a5628b51` Add missing `--cluster {{ cluster }}` on task `set osd_memory_target` in the main.yml file of the ceph-config role. Also it moves the task after ceph configuration file is actually written. Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2022-07-11 19:11:54 +02:00
Guillaume Abrioux	8a5628b516	config/osd: various fixes - sets `osd_memory_target` per osd host. - ceph.conf refactor (osd) Closes: https://bugzilla.redhat.com/show_bug.cgi?id=2056675 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2022-07-11 13:57:32 +02:00
Guillaume Abrioux	5283fa6e96	config: fix indentation in main.yml For consistency and readability. Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2022-07-11 13:57:32 +02:00
Guillaume Abrioux	a99812aa92	facts: follow up on `f6b49f78` `f6b49f78a9` changed a call back to `ipwrap` This fixes this. Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2022-07-06 03:06:01 +02:00
Guillaume Abrioux	45ddbedef2	handler: update ganesha.pid path Due to some changes [1] in nfs-ganesha-4, we now have to use `/var/run/ganesha/ganesha.pid` [1] `52e15c30d0` Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2022-07-05 21:10:50 +02:00
Guillaume Abrioux	434793e2fe	facts: fix set_radosgw_address.yml use `include_tasks` instead of `import_tasks`. Given that with `import_tasks` statements are preprocessed and the tasks that defines it hasn't been run yet, it will fail and complain like following: ``` The error was: 'ansible.vars.hostvars.HostVarsVars object' has no attribute '_interface' ``` Using `include_tasks` instead fixes this. Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2022-07-05 21:10:50 +02:00
Guillaume Abrioux	f6b49f78a9	facts: fix deployments with different net interface names Deployments when radosgws don't have the same names for network interface. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=2095605 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2022-07-05 10:01:40 +02:00
Guillaume Abrioux	2e823b117e	common: fix a typo s/of/or .. Fixes: https://bugzilla.redhat.com/show_bug.cgi?id=2099828#c25 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2022-07-03 07:22:21 +02:00
Guillaume Abrioux	19fedfbac5	nfs: use repo from SIG RPMs for nfs-ganesha aren't hosted anymore at https://download.ceph.com Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2022-06-22 01:17:20 +02:00
Guillaume Abrioux	aa68b06c99	ansible: bump to ansible 2.12 Add required changes to support ansible 2.12 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2022-06-15 08:09:10 +02:00
Michael Wagner	4edaab5f4c	fix(ceph-grafana): make dashboard download work again This fixes the dashboard download for pacific and later. Signed-off-by: Michael Wagner <mitch.wagna@gmail.com>	2022-06-14 14:36:23 +02:00
David Galloway	bcedff95bd	master->main Signed-off-by: David Galloway <dgallowa@redhat.com>	2022-05-30 15:15:15 +02:00
Guillaume Abrioux	c1649862a9	common: move to `ansible.utils.ipwrap` ipwrap has moved to ansible.utils see `db4920ebf6` Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2022-05-12 22:51:31 +02:00
Guillaume Abrioux	1e11f879f6	common: config rhcs tools repo on all nodes Otherwise `cephadm` can't be installed during cephadm-adopt.yml playbook execution. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=2073480 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2022-05-12 22:51:31 +02:00
Ingo Ebel	c5bb450f87	added AlmaLinux and Rocky for iscsi deploy Signed-off-by: Ingo Ebel <ingo.ebel@desy.de>	2022-04-14 00:35:48 +02:00
Guillaume Abrioux	0f34cd16d8	dashboard: allow collecting stats from the host This commit makes podman bindmount `/:/rootfs:ro` so the container can collect data from the host. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=2028775 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2022-04-13 20:35:01 +02:00
pinotelio	f288364c5c	ceph-facts: fix ansible templating error for auto osd discovery This commit fixes templating error that occurs when using auto osd discovery. Getting the len before converting the result to a list causes "object of type generator has no len()" error. Signed-off-by: pinotelio <ahmadreza.mollapour@gmail.com>	2022-04-13 14:26:35 +02:00
Guillaume Abrioux	1cd1fa0560	validate: drop a check Since the ISO install method removal, ceph-ansible isn't able to detect wheter the user is deploying in a 'disconnected environment'. By the way, given that ceph-ansible is available only for upgrading to RHCS 5, this check doesn't make sense anymore, let's drop it. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=2062147 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2022-04-07 13:49:51 +02:00
insatomcat	58fdc03e63	do not update Debian cache when package-install is disabled When deploying with --skip-tags=package-install (when there is no access to a repository), the playbook is still trying to update the package cache, which makes the playbook fail. This change prevents the playbook to try to update the cache when the package-install tag is skipped. Signed-off-by: Florent CARLI <florent.carli@rte-france.com>	2022-03-31 15:28:44 +02:00
Guillaume Abrioux	72e4654aae	dashboard: always set `dashboard_server_addr` When running the playbook with `--limit`, if the play targeted doesn't match hosts present in the mgr group the playbook can fail. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=2063029 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2022-03-25 22:36:23 +01:00
Teoman ONAY	7e8ce2567e	Turn off SELinux separation for containers MON and RGW Initially MONs and RGW binded /etc/pki/ca-trust/extracted using the :z flag (introduced to solve an OSP TripleO issue on RHEL - #3638) but using this flag prevents local services (like sssd) running on the host from accessing the certificates/files in that folder. Signed-off-by: Teoman ONAY <tonay@redhat.com>	2022-03-08 14:45:45 +01:00
Guillaume Abrioux	266b6e739c	adopt: fix node labelling When using group of group, the playbook will apply undesired labels on nodes. This commit fixes it by applying only the expected labels. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=2057528 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2022-03-03 15:52:00 +01:00
Teoman ONAY	da42f3d139	Enable user to change the account used for ssh connection By default cephadm uses root account to connect remotely to other nodes in the cluster. This change allows to choose another account. This commit also allows to use a dedicated subnet for cephadm mgmt. Signed-off-by: Teoman ONAY <tonay@redhat.com>	2022-03-03 15:52:00 +01:00
Seena Fallah	9d87fd87cb	ceph-facts: ignore mounted disks on osd auto discovery Ignore disks with active mountpoint when osd_auto_discovery is true Signed-off-by: Seena Fallah <seenafallah@gmail.com>	2022-02-21 17:15:30 +01:00
Benoît Knecht	7684d892c0	ceph-facts: Fix get_def_crush_rule_name.yml in check mode This construct doesn't work as intended since ansible/ansible#74212: ``` item.stdout \| default('{}') \| from_json ``` That PR made the `command` module return `stdout` even in check mode (setting it to the empty string), so `default()` has no effect in that case and `from_json()` fails to parse an empty string. Instead, `default()` needs to be invoked with its second argument set to `True`, so that it replaces any `False` value (such as an empty string) with its first argument: ``` item.stdout \| default('{}', True) \| from_json ``` Signed-off-by: Benoît Knecht <bknecht@protonmail.ch>	2022-02-07 14:13:19 +01:00
Benoît Knecht	ef05e9a313	ceph-osd: Fix crush_rules.yml in check mode Set a default value for `item.stdout` before passing it to `from_json()`. The `when` condition doesn't prevent this template from being evaluated in check mode, so it fails if `item.stdout` doesn't contain a valid JSON string. Signed-off-by: Benoît Knecht <bknecht@protonmail.ch>	2022-02-07 14:13:19 +01:00
Benoît Knecht	0b3a608216	ceph-osd: Fix start_osds.yml in check mode This construct doesn't work as intended since ansible/ansible#74212: ``` ceph_osd_ids.stdout \| default('{}') \| from_json ``` That PR made the `command` module return `stdout` even in check mode (setting it to the empty string), so `default()` has no effect in that case and `from_json()` fails to parse an empty string. Instead, `default()` needs to be invoked with its second argument set to `True`, so that it replaces any `False` value (such as an empty string) with its first argument: ``` ceph_osd_ids.stdout \| default('{}', True) \| from_json ``` Signed-off-by: Benoît Knecht <bknecht@protonmail.ch>	2022-02-07 14:13:19 +01:00
John Karasev	79ca442d53	ceph-grafana: Add proxy env vars to grafana service template When installing grafana plugins, the container will make http requests. This requires http proxy otherwise installation cannot be performed. Passed the proxy vars from all.yml as env args. Fixes: ceph#6484, ceph#6481 Signed-off-by: John Karasev <john.karasev@intel.com>	2022-02-07 14:09:22 +01:00
Guillaume Abrioux	c491e67486	nfs-ganesha: fix debian based OS deployments Let's use ppa repositories in order to deploy nfs-ganesha on Debian based OS. Fixes: #7031 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2022-01-19 13:42:10 +01:00
Dmitriy Rabotyagov	2eb0a88a67	Use upstream config_template collection In order to reduce need of module internal maintenance and to join forces on plugin development, it's proposed to switch to using upstream version of config_template module. As it's shipped as collection, it's installation for end-users is trivial and aligns with general approach of shipping extra modules. Signed-off-by: Dmitriy Rabotyagov <noonedeadpunk@ya.ru>	2022-01-18 20:22:10 +01:00
Benoît Knecht	bffca06837	ceph-handler: Fix check mode When running in check mode with one or more Ceph daemons that need to be restarted, the `tmpdirpath.path` variable that several handlers rely on is undefined, leading to fatal errors. This commit ensures the tasks that require `tmpdirpath.path` are skipped when it's undefined. Signed-off-by: Benoît Knecht <bknecht@protonmail.ch>	2022-01-06 10:46:55 +01:00
Guillaume Abrioux	dc8940fe1c	common: remove legacy repositories As of rhceph-5, those repositories don't longer exist. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=2032790 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2021-12-15 13:28:50 +01:00
Guillaume Abrioux	f01536ea19	container: align systemd units with rpm Update `After=` and `Wants=` parameters in container systemd units and make them be aligned with the systemd units that come from the packaging. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=2027440 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2021-12-14 13:46:27 +01:00
Danny Webb	189ff93372	make grafana network a configurable option Signed-off-by: Danny Webb <danny.webb@thehutgroup.com>	2021-12-02 08:53:58 +01:00
Guillaume Abrioux	64196ce3a3	validate: support obs repository Otherwise, installation on SuSe fails. Fixes: #6996 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2021-12-02 08:51:45 +01:00
Benoît Knecht	b29a6b18f8	roles/ceph-rgw: Support CRUSH device class The pools created by `ceph-rgw` (listed in `rgw_create_pools`) now support a `ec_crush_device_class` option to specify which device class the EC pool should use. It default to being omitted, which means it will use OSDs from any device class by default. Signed-off-by: Benoît Knecht <bknecht@protonmail.ch>	2021-12-01 08:39:14 +01:00
Guillaume Abrioux	6ad7e52869	validate: fix bug when using vault since a variable encrypted with vault is no longer a string but a encrypted object we can't use the filter \| length, we have to convert it to a string before. Fixes: #6991 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2021-11-16 16:23:27 +01:00
Guillaume Abrioux	82eee4303b	update: support --limit on monitor nodes Change needed in order to support --limit on mon nodes. Otherwise, a call to `hostvars[groups[mon_group_name][0]]['_current_monitor_address']` throws an error: ``` "The error was: 'ansible.vars.hostvars.HostVarsVars object' has no attribute '_current_monitor_address'" ``` Closes: https://bugzilla.redhat.com/show_bug.cgi?id=2014304#c28 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2021-10-28 21:47:01 +02:00
Seena Fallah	4f6da9d92f	ceph-validate: export validate repository vars as a task Signed-off-by: Seena Fallah <seenafallah@gmail.com>	2021-10-08 16:56:47 +02:00
Seena Fallah	e79bda9a05	ceph-common: export repository configuration to a single task Signed-off-by: Seena Fallah <seenafallah@gmail.com>	2021-10-08 16:56:47 +02:00
Seena Fallah	fb99626987	ceph-defaults: set ceph_stable_release default to the stable branch release ceph_stable_release is a legacy from the time where a single branch of ceph-ansible supported more than one release of ceph Signed-off-by: Seena Fallah <seenafallah@gmail.com>	2021-09-30 16:13:55 +02:00
Alex Lambert	a9680ab17f	dashboard: allow disabling of unused features Unconfigured dashboard features can lead to empty tabs in the dashboard containing no meaningful content. Allow users to disable dashboard features they know will not be used. A list of features to be disabled allows the user to define a streamlined dashboard as standard across deployments. Defaults to disabling no features, ensuring that users are sure they do not need the dashboard feature before disabling it. Signed-off-by: Alex Lambert <lamberta@microsoft.com>	2021-09-29 12:02:16 +02:00
Guillaume Abrioux	f8d49827a4	dashboard: retry setting rgw-credentials for some reason, this task can fail in the CI. Adding a retry can help to avoid this failure. Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2021-09-29 10:29:42 +02:00
Guillaume Abrioux	c49d6804bd	common: install ceph-volume package After pacific release, ceph-volume has its own package. ceph-ansible has to explicitly install it on osd nodes. Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2021-09-16 15:12:08 +02:00
Dimitri Savineau	e7b43c1fc6	ceph-defaults: set quay.io as the default registry Because the ceph container images are now only pushed to the quay.io registry then this updates the default registry value. The docker.io registry can still be used but doesn't receive updated container images. Signed-off-by: Dimitri Savineau <dsavinea@redhat.com>	2021-09-09 10:56:09 +02:00
Dimitri Savineau	5bb7240f87	container: explicitly pull monitoring images We don't pull the monitoring container images (alertmanager, prometheus, node-exporter and grafana) in a dedicated task like we're doing for the ceph container image. This means that the container image pull is done during the start of the systemd service. By doing this, pulling the image behind a proxy isn't working with podman. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1995574 Signed-off-by: Dimitri Savineau <dsavinea@redhat.com>	2021-08-23 14:12:45 -04:00
Guillaume Abrioux	6802b8dddd	iscsi: don't set default value for trusted_ip_list It restricts access to the iSCSI API. It can be left empty if the API isn't going to be access from outside the gateway node Even though this seems to be a limited use case, it's better to leave it empty by default than having a meaningless default value. We could make this variable mandatory but that would be a breaking change. Let's just add a logic in the template in order to set this variable in the configuration file only if it was specified by users. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1994930 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> Co-authored-by: Dimitri Savineau <dsavinea@redhat.com>	2021-08-19 09:28:08 -04:00
Guillaume Abrioux	09ef465f62	containers: introduce target systemd unit This adds ceph-*.target systemd unit files support for containerized deployments. This also fixes a regression introduced by PR #6719 (rgw and nfs systemd units not getting purged) Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1962748 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2021-08-18 11:08:50 -04:00
Seena Fallah	95bce32270	ceph-container-engine: allow override container_package_name and container_service_name Only include specific variables when they are undefined Signed-off-by: Seena Fallah <seenafallah@gmail.com>	2021-08-18 09:12:00 +02:00
Guillaume Abrioux	1db8fa8989	roles: remove leftover from pr #4319 pr #4319 introduced some uesless `become: true` on systemd tasks. Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2021-08-18 09:10:15 +02:00
Dimitri Savineau	2ee2194ee0	ceph-dashboard: fix oject gateway integration Since [1] multiple ceph dashboard commands have been removed and this is breaking the current ceph-ansible dashboard with RGW automation. This removes the following dashboard rgw commands: - ceph dashboard set-rgw-api-access-key - ceph dashboard set-rgw-api-secret-key - ceph dashboard set-rgw-api-host - ceph dashboard set-rgw-api-port - ceph dashboard set-rgw-api-scheme Which are replaced by `ceph dashboard set-rgw-credentials` The RGW user creation task is also removed. Finally moving the delegate_to statement from the rgw tasks at the block level. [1] https://github.com/ceph/ceph/pull/42252 Signed-off-by: Dimitri Savineau <dsavinea@redhat.com>	2021-08-17 12:53:58 -04:00
Dimitri Savineau	e44075abd6	ceph-mon: do not log monitor keyring We don't want to display the keyring in the ansible log. Signed-off-by: Dimitri Savineau <dsavinea@redhat.com>	2021-08-12 08:42:05 +02:00
Guillaume Abrioux	7511195738	common: do not log keyring secret let's not display any keyring secret by default in ansible log. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1980744 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2021-08-11 17:33:34 +02:00
Dimitri Savineau	5e0ace7e54	ceph-dashboard: fix TLS cert openssl generation With OpenSSL version prior 1.1.1 (like CentOS 7 with 1.0.2k), the -addext doesn't exist. As a solution, this uses the default openssl.cnf configuration file as a template and add the subjectAltName in the v3_ca section. This temp openssl configuration file is removed after the TLS certificate creation. This patch also move the run_once statement at the block level. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1978869 Signed-off-by: Dimitri Savineau <dsavinea@redhat.com>	2021-08-09 14:19:17 -04:00
Guillaume Abrioux	6f1a0634f7	dashboard: subj_alt_names fact refactor the current way the variable is built results in: ``` 2021-08-03 04:18:23,020 - ceph.ceph - INFO - ok: [ceph-sangadi-4x-indpt6-node1-installer] => changed=false ansible_facts: subj_alt_names: \|- subjectAltName=ceph-sangadi-4x-indpt6-node1-installer/subjectAltName=10.0.210.223/subjectAltName=ceph-sangadi-4x-indpt6-node1-installersubjectAltName=ceph-sangadi-4x-indpt6-node2/subjectAltName=10.0.210.252/subjectAltName=ceph-sangadi-4x-indpt6-node2/ ``` which is incorrect. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1978869 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2021-08-05 18:53:38 -04:00
Teoman ONAY	9b5d97adb9	podman pids.max default value is 2048, docker's one is 4096 which are sufficient for the default value (512) of rgw thread pool size. But if its value is increased near to the pids-limit value, it does not leave place for the other processes to spawn and run within the container and the container crashes. pids-limit set to unlimited regardless of the container engine. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1987041 Signed-off-by: Teoman ONAY <tonay@redhat.com>	2021-08-04 10:20:25 +02:00
Dimitri Savineau	b02cc6931f	ceph-defaults: remove radosgw_civetweb_ variables radosgw_civetweb_xxx variables are legacy variables and users should have switched to radosgw_frontend_xxx variables instead. Signed-off-by: Dimitri Savineau <dsavinea@redhat.com>	2021-08-04 09:13:08 +02:00
Dimitri Savineau	06471a4b82	osds: use osd pool ls instead of osd dump command The ceph osd pool ls detail command is a subset of the ceph osd dump command. $ ceph osd dump --format json\|wc -c 10117 $ ceph osd pool ls detail --format json\|wc -c 4740 Signed-off-by: Dimitri Savineau <dsavinea@redhat.com>	2021-08-02 15:51:01 +02:00
Benoît Knecht	498acd7527	ceph-handler: Fix osd handler in check mode Run the Ceph commands that only gather information (without making any changes to the cluster) when running Ansible in check mode. This allows the tasks that depend on the variables set by those tasks to succeed in check mode. Signed-off-by: Benoît Knecht <bknecht@protonmail.ch>	2021-07-28 14:04:54 +02:00
Dimitri Savineau	f0ccf3ebf0	ceph-defaults: add missing grafana dashboards The radosgw-sync-overview and rbd-details grafana dashboars were missing from the list. Closes: #6758 Signed-off-by: Dimitri Savineau <dsavinea@redhat.com>	2021-07-27 10:49:05 -04:00
Dimitri Savineau	9f77b929d1	alertmanager: allow disable dashboard tls verify When using self-signed/untrusted CA certificates, alertmanager displays an error in logs. With this commit this should make those messages disappear. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1936299 Co-authored-by: Guillaume Abrioux <gabrioux@redhat.com> Signed-off-by: Dimitri Savineau <dsavinea@redhat.com>	2021-07-25 02:56:18 +02:00
Dimitri Savineau	ad05a08160	multisite: use node fqdn for endpoints when https When the rgw_multisite_proto variable is set to https then we shoudn't use the IP address in the zone endpoints list but the node FQDN to match the TLS certificate CN. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1965504 Signed-off-by: Dimitri Savineau <dsavinea@redhat.com>	2021-07-22 21:22:12 +02:00
Dimitri Savineau	cf6e33346e	common: fix py2 pool_list from_json when skipped When using python 2 and the task with a loop is skipped then it generates an error. Unexpected templating type error occurred on ({{ (pool_list.stdout \| from_json)['pools'] }}): expected string or buffer Signed-off-by: Dimitri Savineau <dsavinea@redhat.com>	2021-07-21 08:17:58 +02:00
Guillaume Abrioux	13036115e2	common: disable/enable pg_autoscaler The PG autoscaler can disrupt the PG checks so the idea here is to disable it and re-enable it back after the restart is done. Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2021-07-20 07:37:07 +02:00
Dimitri Savineau	cd06e7c046	ceph-mgr: move mgr module list to common Populating the ceph_mgr_modules list in the mgr_modules doesn't make sense since that file is only executed if the list isn't empty or we're using the dashboard. Signed-off-by: Dimitri Savineau <dsavinea@redhat.com>	2021-07-19 18:23:38 +02:00
Dimitri Savineau	9817d29543	ceph-nfs: allow overriding NFS_CORE_PARAM We already have config override variables for existing block (like ganesha_ceph_export_overrides, ganesha_log_overrides, etc...) or a global one (ganesha_conf_overrides) but redefining the NFS_CORE_PARAM block in that variable will erase all previous values (currently only Bind_Addr). ganesha_core_param_overrides: \| Enable_UDP = false; NFS_Port = 2050; Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1941775 Signed-off-by: Dimitri Savineau <dsavinea@redhat.com>	2021-07-19 18:22:14 +02:00
Guillaume Abrioux	72a0336c71	dashboard: remove "certificate is valid for" error When deploying dashboard with ssl certificates generated by ceph-ansible, we enforce the CN to 'ceph-dashboard' which can makes application such alertmanager complain like following: `err="Post https://mgr0:8443/api/prometheus_receiver: x509: certificate is valid for ceph-dashboard, not mgr0" context_err="context deadline exceeded"` The idea here is to add alternative names matching all mgr/mon instances in the certificate so this error won't appear in logs. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1978869 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2021-07-07 09:38:34 -04:00
Guillaume Abrioux	f4f73b6197	dashboard: support dedicated network for the dashboard This introduces a new variable `dashboard_network` in order to support deploying the dashboard on a different subnet. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1927574 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2021-07-05 21:34:43 +02:00
Dimitri Savineau	1d56818658	prometheus: fix prometheus target url The prometheus service isn't binding on localhost. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1933560 Signed-off-by: Dimitri Savineau <dsavinea@redhat.com>	2021-07-02 17:20:02 +02:00
Dimitri Savineau	d704b05e52	ceph-facts: move device facts to its own file Instead of reusing the condition 'inventory_hostname in groups[osds]' on each device facts tasks then we can move all the tasks into a dedicated file and set the condition on the import_tasks statement. Signed-off-by: Dimitri Savineau <dsavinea@redhat.com>	2021-07-02 14:02:30 +02:00
Dimitri Savineau	55bca07cb6	ceph-validate: check logical volumes We currently don't check if the logical volume used in lvm_volumes list for either bluestore data/db/wal or filestore data/journal exist. We're only doing this on raw devices for batch scenario. Signed-off-by: Dimitri Savineau <dsavinea@redhat.com>	2021-07-02 14:02:30 +02:00
Dimitri Savineau	808e7106de	ceph-validate: check db/journal/wal devices too When using dedicated devices for db/journal/wal objecstore with ceph-volume lvm batch then we should also validate that those devices exist and don't use a gpt partition table in addition of the devices and lvm_volume.data variables. Signed-off-by: Dimitri Savineau <dsavinea@redhat.com>	2021-07-02 14:02:30 +02:00
Dimitri Savineau	7e50380f7f	ceph-validate: use root device from ansible_mounts Instead of using findmnt command to find the device associated to the root mount point then we can use the ansible_mounts fact. Signed-off-by: Dimitri Savineau <dsavinea@redhat.com>	2021-07-02 14:02:30 +02:00
Dimitri Savineau	0df99dda8d	ceph-validate: do not resolve devices This is already done in the ceph-facts role. Signed-off-by: Dimitri Savineau <dsavinea@redhat.com>	2021-07-02 14:02:30 +02:00
Dimitri Savineau	14d458b3b4	ceph-validate: check block presence first Instead of doing two parted calls we can check first if the device exist and then test the partition table. Signed-off-by: Dimitri Savineau <dsavinea@redhat.com>	2021-07-02 14:02:30 +02:00
Dimitri Savineau	ac0342b72e	ceph-validate: check devices from lvm_volumes `2888c08` introduced a regression as the check_devices tasks file was only included based on the devices variable. But that file also validate some devices from the lvm_volumes variable. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1906022 Signed-off-by: Dimitri Savineau <dsavinea@redhat.com>	2021-07-02 14:02:30 +02:00
Dimitri Savineau	9758e3c513	container: set tcmalloc value by default All ceph daemons need to have the TCMALLOC_MAX_TOTAL_THREAD_CACHE_BYTES environment variable set to 128MB by default in container setup. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1970913 Signed-off-by: Dimitri Savineau <dsavinea@redhat.com>	2021-06-30 20:30:55 +02:00
Dimitri Savineau	a05730b38a	rhcs: remove ISO install method Starting RHCS 5, there's no ISO available anymore. This removes all ISO variables and the ceph_repository_type variable. Closes: #6626 Signed-off-by: Dimitri Savineau <dsavinea@redhat.com>	2021-06-30 18:03:03 +02:00
Boris Ranto	2491d4e004	dashboard: Add new prometheus alert It was requested for us to update our alerting definitions to include a slow OSD Ops health check. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1951664 Signed-off-by: Boris Ranto <branto@redhat.com>	2021-06-24 09:02:21 +02:00
Guillaume Abrioux	8279d14d32	multisite: fix bug during switch2containers When running the switch-to-containers playbook with multisite enabled, the fact "rgw_instances" is only set for the node being processed (serial: 1), the consequence of that is that the set_fact of 'rgw_instances_all' can't iterate over all rgw node in order to look up each 'rgw_instances_host'. Adding a condition checking whether hostvars[item]["rgw_instances_host"] is defined fixes this issue. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1967926 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2021-06-17 01:49:29 +02:00
Guillaume Abrioux	8dbee99882	nfs: do no copy client.bootstrap-rgw when using mds There's no need to copy this keyring when using nfs with mds Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2021-06-16 06:32:43 +02:00
Guillaume Abrioux	38bfad46e8	container: conditionnally disable lvmetad Enabling lvmetad in containerized deployments on el7 based OS might cause issues. This commit make it possible to disable this service if needed. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1955040 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2021-06-15 20:16:38 +02:00
Guillaume Abrioux	d58500ade0	ceph_key: handle error in a better way When calling the `ceph_key` module with `state: info`, if the ceph command called fails, the actual error is hidden by the module which makes it pretty difficult to troubleshoot. The current code always states that if rc is not equal to 0 the keyring doesn't exist. `state: info` should always return the actual rc, stdout and stderr. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1964889 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2021-06-14 23:46:20 +02:00
Guillaume Abrioux	f7166cccbf	rolling_update: fix mon+rgw/multisite collocation When monitors and rgw are collocated with multisite enabled, the rolling_update playbook fails because during the workflow, we run some radosgw-admin commands very early on the first mon even though this is the monitor being upgraded, it means the container doesn't exist since it was stopped. This block is relevant only for scaling out rgw daemons or initial deployment. In rolling_update workflow, it is not needed so let's skip it. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1970232 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2021-06-11 10:50:50 +02:00
Neelaksh Singh	d18a9860cd	Sensitive key data now hidden in output log Fixes: #6529 Signed-off-by: Neelaksh Singh <neelaksh48@gmail.com>	2021-06-08 20:46:37 +02:00
Guillaume Abrioux	4daed1f137	dashboard: set cookie_secure in grafana When using grafana behind https `cookie_secure` should be set to `true`. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1966880 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2021-06-04 14:01:28 +02:00
Guillaume Abrioux	664dae0564	prometheus: enforce osd nodes in templates When osd nodes are collocated in the clients group (HCI context for instance), the current logic will exclude osd nodes since they are present in the client group. The best fix would be to exclude clients node only when they are not member of another group but for now, as a workaround, we can enforce the addition of osd nodes to fix this specific case. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1947695 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2021-05-25 16:53:49 +02:00
Guillaume Abrioux	e6d8b058ba	nfs: get org.ganesha.nfsd.conf from container Since we need to revert `33bfb10`, this is an alternative to initial approach. We can avoid maintaining this file since it is present in container image. The idea is to simply get it from the image container and write it to the host. Fixes: #6501 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2021-05-07 13:35:37 +02:00
Dimitri Savineau	a670982a38	ceph-rgw: fix pg_autoscale_mode for pool The pg_autoscale_mode for rgw pools introduced in `9f03a52` was wrong and was missing a `value` keyword because `rgw_create_pools` is a dict. Fixes: #6516 Signed-off-by: Dimitri Savineau <dsavinea@redhat.com>	2021-05-06 10:15:13 +02:00
Guillaume Abrioux	8f87754b76	ceph-nfs: fix dev repo task We need to filter with the OS architecture in order to fetch the right dev repository in shaman Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2021-04-29 19:44:17 +02:00
Seena Fallah	41295f0ef6	ceph-osd: allow to use ceph_tcmalloc_max_total_thread_cache for bluestore TCMALLOC_MAX_TOTAL_THREAD_CACHE_BYTES is for both bluestore and filestore Signed-off-by: Seena Fallah <seenafallah@gmail.com>	2021-04-28 20:03:46 +02:00
Dimitri Savineau	4e6b2a54d2	ceph-defaults: update multisite readme reference The multisite README file has been merged into a single file. Closes: #6411 Signed-off-by: Dimitri Savineau <dsavinea@redhat.com>	2021-04-15 19:44:38 +02:00
Francesco Pantano	441651638d	Config the monitoring stack components api urls using a VIP When dashboard_frontend_vip is provided, all the services should be configured using the related VIP. A new VIP variable is added for both prometheus and alertmanager: we're already able to properly config the grafana vip using dashboard_frontend_vip variable. This change adds the same variable for both prometheus and alertmanager. Signed-off-by: Francesco Pantano <fpantano@redhat.com>	2021-04-15 14:25:53 +02:00
Guillaume Abrioux	839fac8f94	core: bump ansible version We should consider bumping ansible version for future releases, so let's start testing against ansible 2.10 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2021-04-15 13:49:24 +02:00
Benoît Knecht	c078513475	ceph-rgw-loadbalancer: Fix rgw_ports fact The `set_fact rgw_ports` task was failing due to a templating error, because `hostvars[item].rgw_instances` is a list, but it was treated as if it was a dictionary. Another issue was the fact that the `unique` filter only applied to the list being appended to `rgw_ports` instead of the entire list, which means it was possible to have duplicate items. Lastly, `rgw_ports` would have been a list of integers, but the `seport` module expects a list of strings. This commit fixes all of the issues above, allowing the `ceph-rgw-loadbalancer` role to work on systems with SELinux enabled. Signed-off-by: Benoît Knecht <bknecht@protonmail.ch>	2021-04-15 10:39:08 +02:00
Guillaume Abrioux	bab403b603	container/systemd: ensure /var/log/ceph exists This adds a `ExecStartPre=-/usr/bin/mkdir -p /var/log/ceph` in all systemd service templates for all ceph daemon. This is specific to RHCS after a Leapp upgrade is done. Indeed, the `/var/log/ceph` seems to be removed after the upgrade. In order to work around this issue let's ensure the directory is present before trying to start the containers with podman. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1949489 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2021-04-14 16:37:33 +02:00

1 2 3 4 5 ...

2959 Commits (bba9955bf7efdc74cf6d389c29774a9f1664951a)