ceph-ansible

Commit Graph

Author	SHA1	Message	Date
Guillaume Abrioux	709deb90cc	handler: refact check_socket_non_container the `stat --printf=%n` returns something like following: ``` ok: [osd0] => changed=false cmd: \|- stat --printf=%n /var/run/ceph/ceph-osd*.asok delta: '0:00:00.009388' end: '2020-10-06 06:18:28.109500' failed_when_result: false rc: 0 start: '2020-10-06 06:18:28.100112' stderr: '' stderr_lines: <omitted> stdout: /var/run/ceph/ceph-osd.2.asok/var/run/ceph/ceph-osd.5.asok stdout_lines: <omitted> ``` it makes the next task "check if the ceph osd socket is in-use" grep like this: ``` ok: [osd0] => changed=false cmd: - grep - -q - /var/run/ceph/ceph-osd.2.asok/var/run/ceph/ceph-osd.5.asok - /proc/net/unix ``` which will obviously fail because this path never exists. It makes the OSD handler broken. Let's use `find` module instead. Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `46d4d97da9`)	2020-10-14 10:31:05 +02:00
Benoît Knecht	69a6053114	Fix Ansible check mode for site.yml.sample playbook Make sure the `site.yml.sample` playbook can be run in check mode by skipping tasks that try to read the output of commands that have been skipped. Signed-off-by: Benoît Knecht <bknecht@protonmail.ch> (cherry picked from commit `54ba38e35e`)	2020-10-07 07:06:54 +02:00
Guillaume Abrioux	52826caa51	rgw: fix multi instances scaleout in baremetal When rgw and osd are collocated, the current workflow prevents from scaling out the radosgw_num_instances parameter when rerunning the playbook in baremetal deployments. When ceph-osd notifies handlers, it means rgw handlers are triggered too. The issue with this is that they are triggered before the role ceph-rgw is run. In the case a scaleout operation is expected on `radosgw_num_instances` it causes an issue because keyrings haven't been created yet so the new instances won't start. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1881313 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `a802fa2810`)	2020-10-06 09:21:58 -04:00
Seena Fallah	eebed2990d	ceph-facts: add get default crush rule from running monitor In case of deploying new monitor node to an existing cluster, osd_pool_default_crush_rule should be taken from running monitor because ceph-osd role won't be run and the new monitor will have different osd_pool_default_crush_role from other monitors. Signed-off-by: Seena Fallah <seenafallah@gmail.com> (cherry picked from commit `ff9f4d138f`)	2020-09-29 16:38:38 +02:00
Ali Maredia	b753e7db15	rgw multisite: check connection for realm endpoint This commit adds connection checks before realm pulls Curls are performed on the endpoint being pulled from the mons and the rgws Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1731158 Signed-off-by: Ali Maredia <amaredia@redhat.com> (cherry picked from commit `902575369c`)	2020-09-29 16:33:20 +02:00
Dimitri Savineau	fabaec6351	ceph-handler: set handler on xxx_stat result In non containerized deployment we check if the service is running via the socket file presence. This is done via the xxx_socket_stat variable that check the file socket in the /var/run/ceph/ directory. In some scenarios, we could have the socket file still present in that directory but not used by any process. That's why we have the xxx_stat variable which clean those leftovers. The problem here is that we're set the variable for the handlers status (like handler_mon_status) based on xxx_socket_stat instead of xxx_stat. That means we will trigger the handlers if there's an old socket file present on the system without any process associated. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1866834 Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `733596582d`)	2020-09-29 16:33:08 +02:00
Seena Fallah	0dd5036f6c	ceph-facts: check for mon socket in its own host delegate to its own host after checking mon socket to findout if mon socket is in-use or not. Signed-off-by: Seena Fallah <seenafallah@gmail.com> (cherry picked from commit `69f7e35382`)	2020-09-29 16:32:54 +02:00
Guillaume Abrioux	f9a6f775e9	mds: support enabling pg autoscaler on rerun This commit add the pg autoscaler enablement support on ceph-ansible rerun. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1836431 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2020-09-29 16:32:29 +02:00
Dimitri Savineau	7ffd3baa95	ceph-config: remove ceph_release from ceph.conf.j2 We don't use ceph_release variable in the ceph.conf jinja template. Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `62bd41f0d4`)	2020-09-29 16:32:17 +02:00
Dmitriy Rabotyagov	6d5c74aa98	Remove libjemalloc1 installation task libjemalloc1 package is not required neither for ganesha dependency nor for the package build process. So this task can be simply dropped. Signed-off-by: Dmitriy Rabotyagov <noonedeadpunk@ya.ru> (cherry picked from commit `297532ca41`)	2020-09-29 16:30:36 +02:00
Guillaume Abrioux	f9d4eb8b41	facts: refact `ceph_uid` fact There's no need to set this fact with a `set_fact` We can achieve this in `ceph-defaults` Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1875058 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `bcc673f66c`)	2020-09-21 13:49:03 -04:00
Dimitri Savineau	1385d2fdd0	ceph-facts: move facts to defaults value There's no need to define a variable via a fact if we can do it via a default value. Using a fact could be interesseting to override the default value on some condition. - ceph_uid could be set to 167 by default because it's only different on non containerized deployment on Debian/Ubuntu. - rbd_client_directory_{owner,group,mode} could be set to ceph,ceph,0770 by default install of null as we are doing in the facts. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1875058 Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `7f997e623a`)	2020-09-21 13:49:03 -04:00
Dimitri Savineau	9412c44906	container: quote registry password When using a quote in the registry password then we have the following error: The error was: ValueError: No closing quotation To fix this we need to use the quote filter. Close: https://bugzilla.redhat.com/show_bug.cgi?id=1880252 Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `6dcfdf17d4`)	2020-09-18 15:21:32 -04:00
Guillaume Abrioux	1527b9b12a	facts: fix 'set_fact rgw_instances with rgw multisite' the current condition doesn't work, as soon as the first iteration is done the condition makes next iterations skip since `rgw_instances` got set with the first iteration. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1859872 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `ff19c1d851`)	2020-09-18 10:35:28 -04:00
Dimitri Savineau	195ce88e26	ceph-infra: include iscsi nodes for logrotate The iscsi nodes aren't included in the logrotate condition. Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `85643edfe3`)	2020-09-17 14:49:56 -04:00
Guillaume Abrioux	c60a7ad4f6	infra: support log rotation for tcmu-runner This commit adds the log rotation support for tcmu-runner. ceph-container related PR: ceph/ceph-container#1726 Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1873915 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `f576c02ff7`)	2020-09-16 22:37:18 -04:00
Dimitri Savineau	fbc375387a	container: add optional http(s) proxy option When using a http(s) proxy with either docker or podman we can rely on the HTTP_PROXY, HTTPS_PROXY and NO_PROXY environment variables. But with ansible, even if those variables are defined in a source file then they aren't loaded during the container pull/login tasks. This implements the http(s) proxy support with docker/podman. Both implementations are different: 1/ docker doesn't rely en the environment variables with the CLI. Thos are needed by the docker daemon via systemd. 2/ podman uses the environment variables so we need to add them to the login/pull tasks. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1876692 Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `bda3581294`)	2020-09-16 11:32:24 -04:00
Dimitri Savineau	13fb83fc93	ceph-prometheus: update pool stat counter Since [1] The bytes_used pool counter in prometheus has been renamed to stored. Closes: #5781 [1] https://github.com/ceph/ceph/commit/71fe9149 Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `e54b924eaf`)	2020-09-16 10:08:54 -04:00
Dimitri Savineau	fd0b9491b6	ansible: bump to ansible 2.9 Prior this commit we were supporting both ansible 2.8 and 2.9. Let's drop 2.8 now. Closes: #5459 Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1879178 Signed-off-by: Dimitri Savineau <dsavinea@redhat.com>	2020-09-15 13:13:09 -04:00
Dimitri Savineau	5cbbc904c1	node-exporter: exclude client nodes We don't need to install node-exporter on client node because there's no ceph services running on them. This also makes sure we use the group name variables in the prometheus service template instead of hardcoding the values. Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `b105549ed8`)	2020-09-14 16:13:25 -04:00
Guillaume Abrioux	edb7bdd911	Revert "Make 'disable ssl for dashboard task' idempotent." This reverts commit `f607857f2a`. > That commit [1] introduced a regression in the dashboard configuration > because the ceph config get mgr xxxx command doesn't work with > nautilus. > In that release the get operation needs an entity. > [1] `f607857` Signed-off-by: Dimitri Savineau dsavinea@redhat.com	2020-09-11 09:37:23 -04:00
Guillaume Abrioux	44e3195ded	facts: refact and optimize memory consumption there's no need to run this task on all nodes. This uses too much memory for nothing. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1856981 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `f0fe193d8e`)	2020-09-11 09:37:23 -04:00
Guillaume Abrioux	448f36fbbd	config: only add related rgw section there's no need to add each rgw section on all rgw nodes. With this commit, only related rgw section are rendered. Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `0a581a6e60`)	2020-09-10 20:55:07 -04:00
Dimitri Savineau	6177a87185	ceph-iscsi: remove python rtslib shaman repository The rtslib python library is now available in the distribution so we shouldn't have to use the shaman repository Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `254ab54f80`)	2020-09-10 20:38:34 -04:00
Dimitri Savineau	47f24ec047	Add CentOS 8 support for rpm deployment We were only supporting CentOS 8 for containerized deployment. Since Nautilus 14.2.10 we now have el8 rpm packages so we should be able to deploy a nautilus ceph cluster with el8. Note that the nfs-ganesha isn't supported because there's no el8 rpm packages for nfs-ganesha V2.8. Signed-off-by: Dimitri Savineau <dsavinea@redhat.com>	2020-09-10 20:38:34 -04:00
Niko Smeds	67d505af82	Enable HAProxy backend checks for Ceph RGW Add the `check` option to server definitions to enable basic HAProxy health checks for Ceph RADOS gateway backends. Currently traffic will be forwarded to unhealthly `radosgw.service` servers. These changes resolve the issue. Signed-off-by: Niko Smeds nikosmeds@gmail.com (cherry picked from commit `a951c1a3f0`)	2020-09-10 20:38:01 -04:00
Guillaume Abrioux	97a2640714	dashboard: refact admin user creation task this commit splits this task in order to avoid using a `shell` module. Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `54d3e9650f`)	2020-09-10 20:37:42 -04:00
George Shuklin	f607857f2a	Make 'disable ssl for dashboard task' idempotent. This should reduce number of 'changed' tasks during convergence test. Signed-off-by: George Shuklin <george.shuklin@gmail.com> (cherry picked from commit `73d4bb6bd6`)	2020-09-10 20:37:26 -04:00
Rafał Wądołowski	db71eabeef	Comment out ceph_custom_key Since there is a check if ceph_custom_key is defined, there is no reason to define it by default. Signed-off-by: Rafał Wądołowski <rwadolowski@cloudferro.com> (cherry picked from commit `55cd6e83e4`)	2020-09-10 20:37:15 -04:00
Anthony Rusdi	46e4d2aeeb	ceph_custom_repo: define apt and rpm key for custom repo This commit also remove the notify on new added debian repo, force update_cache to yes and define sample ceph_custom_key vars. Signed-off-by: Anthony Rusdi <33247310+antrusd@users.noreply.github.com> (cherry picked from commit `4c592066b7`)	2020-09-10 20:37:15 -04:00
Dimitri Savineau	df70345e6a	ceph-rgw: allow specifying crush rule on pool We already support specifiying a custom crush rule during pool creation in ceph-osd role but not in ceph-rgw role. This patch adds the missing code to implement this feature. Note this is only available for replicated pool not erasure. The rule must also exist prior the pool creation. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1855439 Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `cb8f0237e1`)	2020-09-10 20:36:54 -04:00
Dimitri Savineau	69b09f9336	Allow updating crush rule on existing pool The crush rule value was only set once during the pool creation. It was not possible to update the crush rule value by updating the value in the configuration. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1847166 Signed-off-by: Dimitri Savineau <dsavinea@redhat.com>	2020-09-10 20:35:44 -04:00
Ali Maredia	30d08e1302	rgw: allow rgws to be concurrently with or without multisite Allows rgws in a ceph cluster to be run with multisite and without multisite at the same time. Signed-off-by: Ali Maredia <amaredia@redhat.com> (cherry picked from commit `5c1f4b1a1e`)	2020-09-10 20:35:28 -04:00
Dimitri Savineau	182319d58c	ceph-handler: add missing condition on ceph-crash The ceph-crash tasks present in the ceph-handler role don't need to be executed on all nodes. Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `18e3c7a0a2`)	2020-09-10 20:35:04 -04:00
Guillaume Abrioux	e0ad8194db	crash: rm container in ExecPreStart even with docker We should ensure the container is removed in `ExecPreStart` even when `{{ container_binary }}` is docker. Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `39bb279a53`)	2020-09-10 20:35:04 -04:00
Guillaume Abrioux	66dde0034b	ceph-crash: introduce new role ceph-crash This commit introduces a new role `ceph-crash` in order to deploy everything needed for the ceph-crash daemon. Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `9d2f2108e1`)	2020-09-10 20:35:04 -04:00
Dimitri Savineau	b745c76491	ceph-facts: only get fsid when monitor are present When running the rolling_update playbook with an inventory without monitor nodes defined (like external scenario) then we can't retrieve the cluster fsid from the running monitor. In this scenario we have to pass this information manually (group_vars or host_vars). Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1877426 Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `f63022dfec`)	2020-09-10 17:42:28 -04:00
John Fulton	5b73af9c34	Set default permission for prometheus config files Regardless of the outcome of Ansible 2.9.12 issue 71200 we can set a default permission for these files. Closes: https://github.com/ceph/ceph-ansible/issues/5677 Signed-off-by: John Fulton <fulton@redhat.com> (cherry picked from commit `95dee6f1ca`)	2020-08-18 21:39:56 -04:00
Guillaume Abrioux	03931362dc	mgr: enable pg_autoscaler by default Otherwise, even though we set the pg autoscaler attribute on a pool, the feature won't be working as expected. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1836431 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2020-08-18 14:49:31 -04:00
Guillaume Abrioux	d84161db1a	infra: only install logrotate on right nodes For intsance, there is no need to install logrotate on clients nodes. This also ensure logrotate is installed only for containerized deployments since the packaging has an explicit dependency to logrotate Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `8ed11ea3ee`)	2020-08-18 11:10:38 -04:00
Guillaume Abrioux	51b51b854a	infra: add missing tag This commit adds the missing `with_pkg` tag on the logrotate installation task. Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `e1cb385740`)	2020-08-13 10:09:40 -04:00
Guillaume Abrioux	3987a82d29	infra: add log rotation support (containers) This commit adds the log rotation support via logrotate in containerized deployments. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1848388 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `f1aa6cea21`)	2020-08-13 14:21:44 +02:00
Guillaume Abrioux	88c9f6d969	common: don't enable debug log on ceph-volume calls by default ceph-volume can generate large logs at some point. debug logs by definition should be enabled only when debugging. Let's make it customizable with a variable which is set to `False` by default. Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `448cc280b7`)	2020-08-13 14:21:44 +02:00
Guillaume Abrioux	fa0484d481	nfs: do not copy rgw keyring when `nfs_obj_gw` is true This keyring shouldn't be copied when `nfs_obj_gw` is `True` if the cluster doesn't contain a rgw node, which can be the case given we are using `nfs_obj_gw` instead of `nfs_file_gw` (cephfs vs. object), the deployment will fail trying to copy a key that doesn't exist. Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com> (cherry picked from commit `dd4b5b0328`)	2020-08-12 14:58:13 -04:00
raul	3c1e81ce48	rgw: support 1+ rgw instance in `radosgw_frontend_port` Change the radosgw_frontend_port to take in account more than 1 RGW instance, in it's original form `radosgw_frontend_port: radosgw_frontend_port \| int`, it configured the 8080 port to all instances, with the following modification `radosgw_frontend_port: radosgw_frontend_port \| int + item\|int` we increase in 1 the port count. Co-authored-by: Daniel Parkes <dparkes@redhat.com> Signed-off-by: raul <rmahique@redhat.com> (cherry picked from commit `110eaf5f9f`)	2020-08-12 14:57:44 -04:00
Paulo Matias	102e0337b5	Prometheus APIs are only available through plain http Trying to access these APIs through TLS produces "Could not reach external API" errors in Ceph dashboard. Signed-off-by: Paulo Matias <matias@ufscar.br> (cherry picked from commit `dac8e1d0a9`)	2020-08-06 11:29:25 -04:00
Paulo Matias	ee920b5a9b	Allow user to specify grafana_server_fqdn This is needed to get a TLS certificate to validate correctly. If unspecified, auto-detected grafana_server_addr is used. Signed-off-by: Paulo Matias <matias@ufscar.br> (cherry picked from commit `38ce02c2ea`)	2020-08-06 11:29:25 -04:00
Dimitri Savineau	85dfbc9e0b	dashboard: allow remote TLS cert/key copy When using TLS on the ceph dashboard or grafana services, we can provide the TLS certificate and key. Those files should be present on the ansible controller and they will be copyied to the right node(s). In some situation, the TLS certificate and key could be already present on the target node and not on the ansible controller. For this scenario, we just need to copy the files locally (on each remote host). This patch adds the dashboard_tls_external variable (with default to false) to allow users to achieve this scenario when configuring this variable to true. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1860815 Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `0d0f1e71df`)	2020-08-04 14:02:27 +02:00
Dimitri Savineau	cce042c65b	ceph-handler: remove iscsigws restart scripts The iscsigws restart scripts for tcmu-runner and rbd-target-{api,gw} services only call the systemctl restart command. We don't really need to copy a shell script to do it when we can use the ansible service module instead. Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `cbe79428e6`)	2020-07-27 09:33:00 -04:00
Dimitri Savineau	d408c75d76	podman: always remove container on start In case of failure, the systemd ExecStop isn't executed so the container isn't removed. After a reboot of a failed node, the container doesn't start because the old container is still present in created state. We should always try to remove the container in ExecStartPre for this situation. A normal reboot doesn't trigger this issue and this also doesn't affect nodes running containers via docker. This behaviour was introduced by `d43769d`. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1858865 Signed-off-by: Dimitri Savineau <dsavinea@redhat.com> (cherry picked from commit `47b7c00287`)	2020-07-24 12:47:21 -04:00

1 2 3 4 5 ...

2630 Commits (709deb90cc9a76cc51d784d7e4fd22dda243246f)