ceph-ansible

Commit Graph

Author	SHA1	Message	Date
Guillaume Abrioux	d4e31b90a6	Revert "osd: container remove --pid=host" This reverts commit `bb2bbeb941`. Looks like when not passing `--pid=host` we are facing some issues when deploying more than 2 OSDs in containerized environment. At the moment, we are still troubleshooting this issue but we prefer to revert this commit so it doesn't block any PR in the CI. As soon as we have a fix; we will push a new PR to remove `--pid=host` (a revert of revert...) Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2019-02-14 10:34:37 +00:00
Guillaume Abrioux	500256cdab	validate: fix ntp_daemon_type check in validate is_atomic is defined in ceph-facts or very early in main playbook. In non containerized deployment, is_atomic is only set in ceph-facts which is played after ceph-validate. Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2019-02-14 10:34:37 +00:00
Guillaume Abrioux	76303b457c	container: create ceph-common.conf tmpfiles.d if it doesn't exist Otherwise the task will fail. Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2019-02-14 10:34:37 +00:00
Guillaume Abrioux	b24202f6a4	facts: move two set_fact into ceph-facts those two set_fact tasks should be moved in ceph-facts. Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2019-02-14 10:34:37 +00:00
Guillaume Abrioux	8c8ec63633	container: use tmpfiles.d to creates /run/ceph instead of using `RuntimeDirectory` parameter in systemd unit files, let's use a systemd `tmpfiles.d` to ensure `/run/ceph`. Explanation: `podman` doesn't create the `/var/run/ceph` if it doesn't exist the time where the container is run while `docker` used to create it. In case of `switch_to_containers` scenario, `/run/ceph` gets created by a tmpfiles.d systemd file; when switching to containers, the systemd unit file complains because `/run/ceph` already exists The better fix would be to ensure `/usr/lib/tmpfiles.d/ceph-common.conf` is removed and only rely on `RuntimeDirectory` from systemd unit file parameter but we come from a non-containerized environment which is already running, it means `/run/ceph` is already created and when starting the unit to start the container, systemd will still complain and we can't simply remove the directory if daemons are collocated. Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2019-02-13 09:42:27 +01:00
Guillaume Abrioux	7e0a70f7a8	switch_to_containers: do not try to redeploy monitors `ceph-mon` tries to redeploy monitors because it assumes it was not yet deployed since `mon_socket_stat` and `ceph_mon_container_stat` are undefined (indeed, we stop the daemon before calling `ceph-mon` in the switch_to_containers playbook). Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2019-02-13 09:42:27 +01:00
Rishabh Dave	05ea783eff	fix mistake in task that aborts when ntpd is chosen on Atomic Since it's already confusing whether ntp_daemon_type should be "ntp" or "ntpd", fix the mistake in the title of the task that aborts if ntp_daemon_type is set to "ntpd" and OS being used is Atomic. Signed-off-by: Rishabh Dave <ridave@redhat.com>	2019-02-12 09:09:27 +01:00
Rishabh Dave	bdff3e48fd	don't install NTPd on Atomic Since Atomic doesn't allow any installations and NTPd is not present on Atomic image we are using, abort when ntp_daemon_type is set to ntpd. https://github.com/ceph/ceph-ansible/issues/3572 Signed-off-by: Rishabh Dave <ridave@redhat.com>	2019-02-11 12:02:30 +01:00
Sébastien Han	c69c8c9ac1	mon: do not hardcode ceph uid 167 is the ceph uid for Red Hat based system, thus trying to deploy a monitor on Debian fail since the ceph user id on that system is 64045. This commit uses the ceph_uid variable which contains the right uid based on system/container detection. Closes: https://github.com/ceph/ceph-ansible/issues/3589 Signed-off-by: Sébastien Han <seb@redhat.com>	2019-02-11 09:09:40 +00:00
Leah Neukirchen	4fe7f37849	Fix uses of default(omit) with string concatenation When {{omit}} is concatenated with another string, it expands to something like __omit_place_holder__63eea0d96dd6ed867b95405e11d87dddf61f448d. However, in these use-cases we need an empty string. Regression introduced in `d53f55e807`. Signed-off-by: Leah Neukirchen <leah.neukirchen@mayflower.de>	2019-02-08 16:18:15 +00:00
Patrick C. F. Ernzer	c605ff6a68	setup_ntp: call handler to disable ntpd if chronyd used The task setup chronyd called the handler disable chronyd, which of course defeats the purpose. Changing the task to disable ntpd instead fixes the issue of chronyd being disabled after it got enabled. Fixes: #3582 Signed-off-by: Patrick C. F. Ernzer pcfe@redhat.com	2019-02-08 12:04:44 +01:00
Guillaume Abrioux	d4b3c1d409	iscsi-gws: remove a leftover remove leftover introduced by `9d590f4` Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2019-02-08 01:11:42 +01:00
Guillaume Abrioux	9d590f4339	iscsi: fix permission denied error Typical error: ``` fatal: [iscsi-gw0]: FAILED! => msg: 'an error occurred while trying to read the file ''/home/guits/ceph-ansible/tests/functional/all_daemons/fetch/e5f4ab94-c099-4781-b592-dbd440a9d6f3/iscsi-gateway.key'': [Errno 13] Permission denied: b''/home/guits/ceph-ansible/tests/functional/all_daemons/fetch/e5f4ab94-c099-4781-b592-dbd440a9d6f3/iscsi-gateway.key''' ``` `become: True` is not needed on the following task: `copy crt file(s) to gateway nodes`. Since it's already set in the main playbook (site.yml/site-container.yml) The thing is that the files get generated in the 'fetch_directory' with root user because there is a 'delegate_to' + we run the playbook with `become: True` (from main playbook). The idea here is to create files under ansible user so we can open them later to copy them on the remote machine. Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2019-02-07 17:57:22 +01:00
Sébastien Han	bb2bbeb941	osd: container remove --pid=host Let's try again with the Nautilus release. Closes: https://github.com/ceph/ceph-ansible/issues/1297 Signed-off-by: Sébastien Han <seb@redhat.com>	2019-02-07 12:13:51 +00:00
John Fulton	cc0bf197e1	Fix CNI error when net=host is not used on OSD calls Follow up fix that `410abd7` missed. Related: ceph#3561 Signed-off-by: John Fulton <fulton@redhat.com>	2019-02-05 22:49:01 +00:00
John Fulton	719a25b571	Create Ceph Initial Dirs earlier Include tasks from create_ceph_initial_dirs earlier during ceph config role. Fixes: #3568 Signed-off-by: John Fulton <fulton@redhat.com>	2019-02-05 18:38:05 +00:00
John Fulton	dab3f6ee3f	Fix CNI error when net=host is not used in some podman calls With 'podman version 1.0.0' on RHEL8 beta the 'get ceph version' and 'ceph monitor mkfs' commands fail [1] with "error configuring network namespace for container Missing CNI default network". When net=host is added these errors are resolved. net=host is used in many other calls (grep -R net=host \| wc -l --> 38). Fixes: #3561 Signed-off-by: John Fulton <fulton@redhat.com> (cherry picked from commit `410abd7745`)	2019-02-05 18:14:28 +01:00
Guillaume Abrioux	914d94cae8	set RuntimeDirectory in all systemd unit templates /var/run/ceph resides in a non persistent filesystem (tmpfs) After a reboot, all daemons won't start because this directory will be missing. Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2019-02-05 18:14:28 +01:00
Guillaume Abrioux	7ade032807	osd: bind mount /var/run/udev/ without this, the command `ceph-volume lvm list --format json` hangs and takes a very long time to complete. Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2019-02-05 18:14:28 +01:00
Guillaume Abrioux	fdca29f2a7	facts: set timeout_command fact in ceph-defaults - also add `--foreground` which seems to fix some issue we are facing when using timeout with `podman`. - use this fact in the `is ceph running already?` task. Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2019-02-05 18:14:28 +01:00
Guillaume Abrioux	16efdbc59b	podman: support podman installation on rhel8 Add required changes to support podman on rhel8 Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1667101 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2019-02-05 18:14:28 +01:00
John Fulton	37b5d1084a	Make python print statements python3 compatible The restart_osd_daemon.sh generated from the j2 template contains a python call which uses 'print x' instead of 'print(x)'. Add the missing parentheses to make this call compatible with both 2 and 3. Also add parentheses to other python print calls found in roles/ceph-client/defaults/main.yml and infrastructure-playbooks/cluster-os-migration.yml. Fixes: https://bugzilla.redhat.com/show_bug.cgi?id=1671721 Signed-off-by: John Fulton <fulton@redhat.com>	2019-02-01 15:23:27 +00:00
Andrew Schoen	70a4368bc5	ceph-config: do not always assume containers when calculating num_osds CEPH_CONTAINER_IMAGE should be None if containerized_deployment is False. Signed-off-by: Andrew Schoen <aschoen@redhat.com>	2019-02-01 12:28:12 +01:00
Andrew Schoen	88eda479a9	ceph-facts: generate devices when osd_auto_discovery is true This task used to live in ceph-osd, but we need it defined here to that ceph-config can use it when trying to determine the number of osds. Signed-off-by: Andrew Schoen <aschoen@redhat.com>	2019-02-01 12:28:12 +01:00
John Fulton	cba9b23363	Do not timeout podman/docker pull if timeout value is '0' If user sets "docker_pull_timeout: '0'" then do not use the timeout command when running podman/docker pull. Also, use "timeout -s KILL"; without KILL, podman on RHEL8 beta does not timeout and deployment can hang. Related: https://bugzilla.redhat.com/show_bug.cgi?id=1670625 Signed-off-by: John Fulton <fulton@redhat.com>	2019-01-31 15:35:09 +00:00
Guillaume Abrioux	fe1528adb4	config: support num_osds fact setting in containerized deployment This part of the code must be supported in containerized deployment Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1664112 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2019-01-31 14:30:23 +00:00
Ramana Raja	dfff89ce67	Install nfs-ganesha stable v2.7 nfs-ganesha v2.5 and 2.6 have hit EOL. Install nfs-ganesha v2.7 stable that is currently being maintained. Signed-off-by: Ramana Raja <rraja@redhat.com>	2019-01-30 14:57:26 +01:00
Guillaume Abrioux	9f16501747	common: clean monitor initial keyring code let's not be blocked by the fact we don't have the initial keyring in `{{ fetch_directory }}` Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2019-01-30 10:36:02 +01:00
Guillaume Abrioux	f8aa8cdf60	facts: clean fsid generation code clean some leftover and duplicate code. Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2019-01-30 10:36:02 +01:00
Patrick Donnelly	080ce7dd72	do not set ceph_release to dummy When we're deploying a dev branch. Signed-off-by: Patrick Donnelly <pdonnell@redhat.com>	2019-01-29 16:04:44 +00:00
Zack Cerza	82897c76fb	Fix ceph.conf generation `877979c78` in #3486 broke ceph.conf generation entirely. Remove the stray curly brace. Signed-off-by: Zack Cerza <zack@redhat.com>	2019-01-25 09:19:24 +01:00
Guillaume Abrioux	0bfefdd5bc	override ceph_release with ceph_stable_release when `ceph_origin` is set to `'repository'` and `ceph_repository` to `'community'` we need to ensure `ceph_release` reflect `ceph_stable_release`. `4a3f180f9d` simply removed the override while it should just have to be run only when the condition mentioned above is satisfied. Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2019-01-22 15:53:32 +01:00
Guillaume Abrioux	877979c787	config: make sure ceph_release is set for all client node `ceph_release` is set in `ceph-container-common` but this role is played only on first node for clients, this means ceph-config will fail on all client nodes except the first one. This commit ensure ceph_release is set for all client nodes. Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2019-01-22 13:45:38 +01:00
Sébastien Han	5babc1b4eb	mon: enable msgr2 Enabling msgr2 style declaration for Nautilus and above. Prior releases will keep the right syntax. When upgrading from Mimic to Nautilus we must maintain something in the form of: mon_host = [v1:127.0.0.1:6789/0,v2:127.0.0.1:3300/0] Signed-off-by: Sébastien Han <seb@redhat.com>	2019-01-22 13:45:38 +01:00
Sébastien Han	fc34fb1bd9	mon: ability to change mon listening port on container You can now use 'ceph_mon_container_listen_port' to change the port the monitor will listen on. Setting the default to 3300 (assigned by IANA) since Nautilus has released the messenger2 transport protocol. Signed-off-by: Sébastien Han <seb@redhat.com>	2019-01-22 13:45:38 +01:00
Sébastien Han	41a7cc878c	Revert "mon: force peer addition" This reverts commit `ee08d1f89a` which was mostly to workaround a bug in ceph@master. Now, ceph@master is fixed so reverting this. Thanks to https://github.com/ceph/ceph/pull/25900 Signed-off-by: Sébastien Han <seb@redhat.com>	2019-01-22 13:45:38 +01:00
guihecheng	1ac94c048f	rgw: add support for multiple rgw instances on a single host With this, we could have multiple rgw instances on a single host with a single run, don't have to use rgw-standalone.yml which does not seems able to bind ports separately. If you want to have multiple rgw instances, just change 'radosgw_instances' to the number you want, which defaults to 1. Not compatible with Multi-Site yet. Signed-off-by: guihecheng <guihecheng@cmiot.chinamobile.com>	2019-01-18 11:12:28 +01:00
Guillaume Abrioux	1bbdde272f	config: remove code related to ceph release prior to luminous This part of the code is not needed since ceph-ansible@master is intended to deploy ceph@master only. Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2019-01-14 14:41:13 +00:00
Guillaume Abrioux	e9188cd202	ceph-default: rm useless condition This condition is useless and it's also creating issues we don't see in our CI. ceph_release is set by either ceph-common or ceph-docker-common so let's keep it this way. Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1645379 Signed-off-by: Guillaume Abrioux <gabrioux@redhat.com>	2019-01-14 14:41:13 +00:00
Sébastien Han	ee08d1f89a	mon: force peer addition Somewhat something changed with the introduction of msg2 and we have to add each node as a peer so the monitors can form a quorum. This might be due to our CI environment, although adding this is completly harmless and solves monitors not being able to form quorum. It seems that the initial monitor map wasn't containing the right information about the peers (addresses like 0.0.0.0/0r1, for each rank. Signed-off-by: Sébastien Han <seb@redhat.com>	2019-01-09 13:15:52 +01:00
Bruceforce	446f3c9fae	nfs-ganesha: fixed nfs_ganesha_dev_apt_repo variable The nfs_ganesha_dev_apt_repo variable was set incorrect in task "fetch nfs-ganesha development repository" Signed-off-by: Bruceforce <Bruceforce@users.noreply.github.com>	2019-01-05 16:04:05 +01:00
Sébastien Han	9abf9dba0b	rgw: do not create mandatory directories The packages are responsible for this, currently tracked in ceph https://github.com/ceph/ceph/pull/25503 Signed-off-by: Sébastien Han <seb@redhat.com>	2019-01-04 13:57:40 +00:00
Sébastien Han	2af624dc5b	rbd-mirror: copy bootstrap key after package install If we don't copy the key after the package install the directory /var/lib/ceph/bootstrap-rbd-mirror will not exist and the copy will fail. Signed-off-by: Sébastien Han <seb@redhat.com>	2019-01-04 13:57:40 +00:00
Sébastien Han	b1dfe3f03e	config: only pre-create ceph dirs on containers We don't need to create the directories on non-containers, they are created by the packages. Closes: https://github.com/ceph/ceph-ansible/issues/3430 Signed-off-by: Sébastien Han <seb@redhat.com>	2019-01-04 13:57:40 +00:00
Rishabh Dave	6fa757d343	ceph-infra: disable unrequired NTP services When one of the currently supported NTP services has been set up, disable rest of the NTP services on Ceph nodes. Signed-off-by: Rishabh Dave <ridave@redhat.com>	2019-01-04 14:01:05 +01:00
Rishabh Dave	b03ab60742	ceph-infra: merge ntp_debian.yml and ntp_rpm.yml Merge ntp_debian.yml and ntp_rpm.yml into one (the new file is called setup_ntp.yml) since they are almost identical. Also avoid repetition of the common setup step for ntpd and chronyd services. Signed-off-by: Rishabh Dave <ridave@redhat.com>	2019-01-04 14:01:05 +01:00
Rishabh Dave	a14bfa282a	copy certificates as root user Since the current user on the controller node, might not have the permission to read the TLS certificate and related files, copy these files to the Ceph nodes as root user. Fixes: https://github.com/ceph/ceph-ansible/issues/3465 Signed-off-by: Rishabh Dave <ridave@redhat.com>	2019-01-03 09:49:06 +01:00
Kai Wembacher	1dd26f76bf	document missing support for non-containerized deployment Signed-off-by: Kai Wembacher <kai@ktwe.de>	2018-12-21 15:37:55 +00:00
jtudelag	23ad5fd9cb	Clarify RGWs configuration when using ceph_conf_overrides. To avoid future misconfigurations, clarify that the only valid scheme is [client.rgw.] instead of [client.radosgw.].	2018-12-20 13:55:03 +00:00
Kai Wembacher	a273ed7f60	add support for rocksdb and wal on the same partition in non-collocated Signed-off-by: Kai Wembacher <kai@ktwe.de>	2018-12-20 14:19:46 +01:00

1 2 3 4 5 ...

2131 Commits (a1a871cadee5e86d181e1306c985e620b81fccac)