Skip to content

Instantly share code, notes, and snippets.

@bonyiii
Created April 19, 2011 06:54
Show Gist options
  • Select an option

  • Save bonyiii/926930 to your computer and use it in GitHub Desktop.

Select an option

Save bonyiii/926930 to your computer and use it in GitHub Desktop.
corosync pacemaker in openvz

According this thread setting up /dev/shm solves these kind of problems:

Apr 19 06:18:27 mongo2.example.com crmd: [1938]: info: crm_timer_popped: Wait Timer (I_NULL) just popped!
Apr 19 06:18:28 mongo2.example.com crmd: [1938]: info: do_cib_control: Could not connect to the CIB service: connection failed
Apr 19 06:18:28 mongo2.example.com crmd: [1938]: WARN: do_cib_control: Couldn't complete CIB registration 29 times... pause and retry
Apr 19 06:18:30 mongo2.example.com crmd: [1938]: info: crm_timer_popped: Wait Timer (I_NULL) just popped!
Apr 19 06:18:31 mongo2.example.com crmd: [1938]: info: do_cib_control: Could not connect to the CIB service: connection failed
Apr 19 06:18:31 mongo2.example.com crmd: [1938]: WARN: do_cib_control: Couldn't complete CIB registration 30 times... pause and retry
Apr 19 06:18:31 mongo2.example.com crmd: [1938]: ERROR: do_cib_control: Could not complete CIB registration  30 times... hard error
Apr 19 06:18:31 mongo2.example.com crmd: [1938]: ERROR: do_log: FSA: Input I_ERROR from do_cib_control() received in state S_STARTING
Apr 19 06:18:31 mongo2.example.com crmd: [1938]: info: do_state_transition: State transition S_STARTING -> S_RECOVERY [ input=I_ERROR cause=C_FSA_INTERNAL origin=do_cib_control ]
Apr 19 06:18:31 mongo2.example.com crmd: [1938]: ERROR: do_recover: Action A_RECOVER (0000000001000000) not supported
Apr 19 06:18:31 mongo2.example.com crmd: [1938]: info: crm_cluster_connect: Connecting to OpenAIS
Apr 19 06:18:31 mongo2.example.com crmd: [1938]: info: init_ais_connection_once: Creating connection to our AIS plugin
Apr 19 06:18:31 mongo2.example.com crmd: [1938]: info: init_ais_connection_once: Connection to our AIS plugin (9) failed: Library error (2)
Apr 19 06:18:31 mongo2.example.com crmd: [1938]: ERROR: do_log: FSA: Input I_ERROR from do_ha_control() received in state S_RECOVERY
Apr 19 06:18:31 mongo2.example.com crmd: [1938]: info: do_dc_release: DC role released
Apr 19 06:18:31 mongo2.example.com crmd: [1938]: info: do_te_control: Transitioner is now inactive
Apr 19 06:18:31 mongo2.example.com crmd: [1938]: ERROR: do_started: Start cancelled... S_RECOVERY
Apr 19 06:18:31 mongo2.example.com crmd: [1938]: ERROR: do_log: FSA: Input I_TERMINATE from do_recover() received in state S_RECOVERY
Apr 19 06:18:31 mongo2.example.com crmd: [1938]: info: do_state_transition: State transition S_RECOVERY -> S_TERMINATE [ input=I_TERMINATE cause=C_FSA_INTERNAL origin=do_recover ]
Apr 19 06:18:31 mongo2.example.com crmd: [1938]: info: do_lrm_control: Disconnected from the LRM
Apr 19 06:18:31 mongo2.example.com crmd: [1938]: info: do_ha_control: Disconnected from OpenAIS
Apr 19 06:18:31 mongo2.example.com crmd: [1938]: info: do_cib_control: Disconnecting CIB
Apr 19 06:18:31 mongo2.example.com crmd: [1938]: info: do_exit: Performing A_EXIT_0 - gracefully exiting the CRMd
Apr 19 06:18:31 mongo2.example.com crmd: [1938]: ERROR: do_exit: Could not recover from internal error
Apr 19 06:18:31 mongo2.example.com crmd: [1938]: info: free_mem: Dropping I_RELEASE_SUCCESS: [ state=S_TERMINATE cause=C_FSA_INTERNAL origin=do_dc_release ]
Apr 19 06:18:31 mongo2.example.com crmd: [1938]: info: free_mem: Dropping I_TERMINATE: [ state=S_TERMINATE cause=C_FSA_INTERNAL origin=do_stop ]
Apr 19 06:18:31 mongo2.example.com crmd: [1938]: info: do_exit: [crmd] stopped (2)
Apr 19 06:18:31 corosync [pcmk  ] ERROR: pcmk_wait_dispatch: Child process crmd exited (pid=1938, rc=2)
Apr 19 06:18:31 corosync [pcmk  ] notice: pcmk_wait_dispatch: Respawning failed child process: crmd
Apr 19 06:18:31 corosync [pcmk  ] info: spawn_child: Forked child 1947 for process crmd
Apr 19 06:18:31 mongo2.example.com crmd: [1947]: info: Invoked: /usr/lib/heartbeat/crmd
Apr 19 06:18:31 mongo2.example.com crmd: [1947]: info: main: CRM Hg Version: da7075976b5ff0bee71074385f8fd02f296ec8a3
Apr 19 06:31:53 mongo2.example.com crmd: [1958]: info: crmd_init: Starting crmd
Apr 19 06:31:54 mongo2.example.com crmd: [1958]: info: do_cib_control: Could not connect to the CIB service: connection failed
Apr 19 06:31:54 mongo2.example.com crmd: [1958]: WARN: do_cib_control: Couldn't complete CIB registration 1 times... pause and retry
Apr 19 06:31:54 mongo2.example.com crmd: [1958]: info: crmd_init: Starting crmd's mainloop
Apr 19 06:31:56 mongo2.example.com crmd: [1958]: info: crm_timer_popped: Wait Timer (I_NULL) just popped!
Apr 19 06:31:57 mongo2.example.com crmd: [1958]: info: do_cib_control: Could not connect to the CIB service: connection failed
Apr 19 06:31:57 mongo2.example.com crmd: [1958]: WARN: do_cib_control: Couldn't complete CIB registration 2 times... pause and retry
Apr 19 06:31:59 mongo2.example.com crmd: [1958]: info: crm_timer_popped: Wait Timer (I_NULL) just popped!
Apr 19 06:32:00 mongo2.example.com crmd: [1958]: info: do_cib_control: Could not connect to the CIB service: connection failed
Apr 19 06:32:00 mongo2.example.com crmd: [1958]: WARN: do_cib_control: Couldn't complete CIB registration 3 times... pause and retry
Apr 19 06:32:02 mongo2.example.com crmd: [1958]: info: crm_timer_popped: Wait Timer (I_NULL) just popped!

So basically only thing needs to be done in order to run pacemaker in OpenVZ container is to run this command in the VE:

mount -t tmpfs tmpfs /dev/shm/

Delete resource

Add resource first

crm configure primitive ClusterIP ocf:heartbeat:IPaddr2 \
params ip=192.168.45.253 cidr_netmask=32 \
op monitor interval=30s

And for example delete ClusterIP address

crm_resource -D -t primitive -r <id_of_the_resource>

crm_resource -D -t primitive -r ClusterIP

Apache setup

Following vAntMet's blog post. Change in

/usr/lib/ocf/resource.d/heartbeat/apache

from

HTTPDOPTS="-DSTATUS"

to

HTTPDOPTS="-D DEFAULT_VHOST -D INFO -D SSL -D SSL_DEFAULT_VHOST -D LANGUAGE -D STATUS"

or whatever you have in /etc/conf.d/apache2

Apache Debugging

After that one can try out if the config is working by

export OCF_ROOT=/urs/lib/ocf

and run

/usr/lib/ocf/resource.d/heartbeat/apache start

unset OCF_ROOT afteward

 unset OCF_ROOT

Default server-status is query on port 443 and when wget queries the server status server says:

<!DOCTYPE HTML PUBLIC "-//IETF//DTD HTML 2.0//EN">
<html><head>
<title>400 Bad Request</title>
</head><body>
<h1>Bad Request</h1>
<p>Your browser sent a request that this server could not understand.<br />
Reason: You're speaking plain HTTP to an SSL-enabled server port.<br />
Instead use the HTTPS scheme to access this URL, please.<br />
<blockquote>Hint: <a href="https://localhost/"><b>https://localhost/</b></a></blockquote></p>
<hr>
<address>Apache Server at localhost Port 443</address>
</body></html>

During debugging in apache script file one can set

OCF_RESKEY_statusurl="http://127.0.0.1/server-status"

and errors are gone

To make this achivement permanent set up resource this way

crm configure primitive WebSite ocf:heartbeat:apache params \
configfile=/etc/apache2/httpd.conf statusurl=http://127.0.0.1/server-status \ 
op monitor interval=1min

STONITH SSH setup

Create a file configure.txt with the following content

configure
primitive st-ssh stonith:external/ssh \
params hostlist="mongo1.raven.maut mongo2.raven.maut"
clone fencing st-ssh
commit

and run

crm  < configure.txt

clone and primitive are separate commands but somehow needs to be run in one batch.

Here are some other examples

STONITH debug

Stonith scipts are located at: /usr/lib/stonith/plugins

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment