# Far-Edge Subcloud Installation and Remediation Guide ## Installation A Wind River subcloud installation normally runs with no manual intervention. In the event that an installation needs to be started manually: - Visit https://middleware.vcpfe.vzwops.com/caas/wrfix/ in a browser. - Enter the iLO IP address into the first text field (even though it's labeled "cluster name"). - Type ```wr``` in the playbook key field. - Click the submit button. If an OK message appears, the job has been successfully sent to ansible-queue. The remaining steps require an SSH shell on ```root@vcpfe-queue-2-prod-birmingham```: - View ```/var/log/lisp/ansible-queue.log``` to ensure that the playbook started. - View the specific ansible log in ```/var/log/ansible/queue/``` to watch the playbook's progress. If you want to watch the OS installation (the first stage of the overall deployment), attach to the iLO console: - Find the subcloud data (see below for the query) and find the ```ilo_host_address``` and ```bmc_password```. - ```# ssh bladecenter-lb-vip.meter.vzwops.com``` - ```# do-bmc ``` - Enter the password when prompted. - Type ```vsp``` to start the serial console. ## Remediation All too often it is necessary to remediate failed Wind River installations. Even successful ones need a visit by an engineer to ensure that they are in fact complete. It is not possible to discover this information from the middleware database, the subcloud, nor the central controller alone. All three of these data sources must be accessed. TODO: write a dashboard which aggregates the required information from the central controller, subcloud and database. Common causes of failure (this is not an exhaustive list): - OAM network unreachable. - MGMT network unreachable. - MGMT subnet overlaps with another already-installed subcloud. - The subcloud does not sync in time after coming online. - The platform-integ-apps application fails to apply. (This does not include servers that have issues which prevent them from advancing through the automation pipeline far enough to begin a Wind River installation. These include, but are not limited to, hardware problems and incorrect firmware versions.) ### Inspecting a subcloud Three data sources were listed above. Let's visit each one in detail. #### Middleware database The middlware database is our record of what should be installed and whether or not we have marked it as complete. The ```caas_wrbatch``` view and ```caas_wrinstallschedule``` table are important in this context. Here is a handy query to pull up all relevant information by cluster name: select b.cluster_name || '.vcpfe.vzwops.com' as cluster_name, b.*, s.* from caas_wrbatch b left outer join caas_wrinstallschedule s on b.cluster_id = s.cluster_id inner join caas_location l on b.fuze_spm_site_id = l.fuze_spm_site_id where b.cluster_name in ('') order by b.parent_cluster_name asc, b.cluster_name asc; (The expression in the first column is a convenience: it builds a FQDN for the subcloud that can be directly used in an SSH command from your laptop.) Here is a sample record showing a completed subcloud: cluster_name | wsbomagj-d663366-001.vcpfe.vzwops.com id | 423 cluster_id | 423 fuze_spm_site_id | 663366 parent_oam_vip_hostname | parent_oam_vip_address | 2001:4888:a12:3221:106:290:0:10 parent_mgmt_address_range_start | 2001:4888:a12:3220:106:290:0:10 parent_mgmt_address_range_end | 2001:4888:a12:3220:106:290:0:ffff parent_mgmt_default_gateway | 2001:4888:a12:3220:106:28:: parent_mgmt_subnet | 2001:4888:a12:3220:: pxe_mac_address | 48:df:37:e2:67:34 ilo_hostname | wsbomagj-663366-rh-pe0e910-001 ilo_host_address | 2001:4888:2a10:90fd:103:40a:0:e001 oam_hostname | wsbomagj-663366-rh-pe2e910-001 oam_host_address | 2001:4888:2a10:90fd:103:40a:0:400 oam_vip_address | 2001:4888:2a10:90fd:103:40a:0:f400 oam_default_gateway | 2001:4888:2a10:90fd:103:2a0:: mgmt_address_range_start | 2001:4888:2a10:982b:103:40a:: mgmt_address_range_end | 2001:4888:2a10:982b:103:40a:0:f mgmt_default_gateway | 2001:4888:2a10:982b:103:2a0:: mgmt_subnet | 2001:4888:2a10:982b:103:40a:: host_vlan | 2845 oam_vlan | 410 mgmt_vlan | 420 cluster_name | wsbomagj-d663366-001 parent_cluster_name | wsbomagj-c319918-003 vendor_name | HPE intel_nic_firmware_version | 1.2585.0 maint_window_p | f namespace_id | 425 namespace_name | wsbomagj-663366vzwcvdu-y-ss-x-05690012222 bmc_username | k8sctl bmc_password | id | 709 date_scheduled | 2020-11-19 19:15:16.885864+00 date_completed | 2020-12-04 19:28:17.644994+00 date_last_failed | 2020-12-03 20:47:18.199344+00 kirke_ticket_number | 709 kirke_ticket_status | kirke_ticket_completed | 2020-11-20 02:28:30.908054+00 cluster_id | 423 The key points: - ```pxe_mac_address``` is not null - ```intel_nic_firmware_version``` == 1.2585.0 - ```date_completed``` is not null - ```kirke_ticket_number``` is not null The last point is the crucial one. Setting ```kirke_ticket_number``` to something other than null is the way we manually mark the Wind River installation as complete. It must be done with a query similar to the following: update caas_wrinstallschedule set kirke_ticket_number = id where id = 709; #### Central controller Use the following command to see the status of every subcloud attached to a central controller: dcmanager subcloud list If the subcloud in question is complete and healthy from the central controller's point of view, it will look like the following. This does not necessarily mean that all is well. The subcloud itself must be visited and checked. [sysadmin@controller-0 ~(keystone_admin)]$ dcmanager subcloud list | grep wsbomagj-d663366-001 | 546 | wsbomagj-d663366-001 | managed | online | complete | in-sync | #### Subcloud If the central controller says that the subcloud is offline, the first thing to check is the mgmt network. If you cannot ping the central controller's mgmt IP, there is no hope for remediating the subcloud. ping6 If the subcloud is managed, online and in-sync, then the remaining checks can be performed. Check the alarms: fm alarm-list Check the applied status of the applications: [sysadmin@controller-0 ~(keystone_admin)]$ system application-list +--------------------------+----------+-----------------------------------+------------------------------------+---------+-----------+ | application | version | manifest name | manifest file | status | progress | +--------------------------+----------+-----------------------------------+------------------------------------+---------+-----------+ | cert-manager | 20.06-4 | cert-manager-manifest | certmanager-manifest.yaml | applied | completed | | nginx-ingress-controller | 20.06-0 | nginx-ingress-controller-manifest | nginx_ingress_controller_manifest. | applied | completed | | | | | yaml | | | | | | | | | | | oidc-auth-apps | 20.06-26 | oidc-auth-manifest | manifest.yaml | applied | completed | | platform-integ-apps | 20.06-9 | platform-integration-manifest | manifest.yaml | applied | completed | +--------------------------+----------+-----------------------------------+------------------------------------+---------+-----------+ Check the drbd resize. drbd8 should be set to 32GB: [sysadmin@controller-0 ~(keystone_admin)]$ lsblk | grep drbd8 │ └─drbd8 147:8 0 32G 0 disk /var/lib/docker-distribution Check the PTP configuration. ```tx_timestamp_timeout``` should be set to 50: [sysadmin@controller-0 ~(keystone_admin)]$ grep tx_timestamp_timeout /etc/ptp4l.conf tx_timestamp_timeout 50 Check the ISOL CPU configuration which should be enabled on cores 3, 4, 27 and 28: [sysadmin@controller-0 ~(keystone_admin)]$ system host-cpu-list controller-0 | grep -F 'Application-isolated' | 72499a12-e2f2-40fc-a069-a2b002908843 | 3 | 0 | 3 | 0 | Intel(R) Xeon(R) Gold 6212U CPU @ 2.40GHz | Application-isolated | | 8498669d-9048-45f6-b9dd-c81262934b70 | 4 | 0 | 4 | 0 | Intel(R) Xeon(R) Gold 6212U CPU @ 2.40GHz | Application-isolated | | 4a70893f-b2cf-4405-9b69-11ed823a81b3 | 27 | 0 | 3 | 1 | Intel(R) Xeon(R) Gold 6212U CPU @ 2.40GHz | Application-isolated | | 1d6e330d-7d2f-416c-b375-9b909d6a1f3c | 28 | 0 | 4 | 1 | Intel(R) Xeon(R) Gold 6212U CPU @ 2.40GHz | Application-isolated | ### Marking an installation as complete If the subcloud passes all of the checks, it should be marked complete with the update query: update caas_wrinstallschedule set kirke_ticket_number = id where id = ; ### Remediating If the subcloud fails to pass all of the checks, it must be remediated. Deciding which steps to take requires operational experience. Here is a guide to some common problems and solutions: - The subcloud is unmanaged: ping the central controller mgmt IP. If it fails, the network is misconfigured and must be fixed by the region. - ```platform-integ-apps``` is not applied: perform a wipedisk and reinstall. - ```oidc-auth-apps``` is uploaded but not applied: run the ```wr_remediate``` playbook. - The drbd resize hasn't happened: run the ```wr_remediate``` playbook. - PTP isn't configured: run the ```wr_remediate``` playbook. - ISOL CPUS are not configured: run the ```wr_remediate``` playbook. The very last thing that the ```wr-installer``` playbook does is a ```collect all```. You can tell if the playbook ran to completion by checking the contents of the ```/scratch``` directory on the subcloud. If it contains a collect tarball, the playbook completed.