summaryrefslogtreecommitdiff
path: root/doc/markdown/subcloud-install-remediation.md
diff options
context:
space:
mode:
authorckonstanski <kostcarl@isu.edu>2026-07-30 18:17:14 -0600
committerckonstanski <kostcarl@isu.edu>2026-07-30 18:17:14 -0600
commit640ff61422bc0ee3966941b93d9889ddbbd38d77 (patch)
treeabf1c08f5d38ee53ec8b29dc4f425722517505e6 /doc/markdown/subcloud-install-remediation.md
parent22dae02a86c1fce71091bfa5289cf99abba1b217 (diff)
more filesHEADmaster
Diffstat (limited to 'doc/markdown/subcloud-install-remediation.md')
-rw-r--r--doc/markdown/subcloud-install-remediation.md253
1 files changed, 253 insertions, 0 deletions
diff --git a/doc/markdown/subcloud-install-remediation.md b/doc/markdown/subcloud-install-remediation.md
new file mode 100644
index 0000000..753c5fb
--- /dev/null
+++ b/doc/markdown/subcloud-install-remediation.md
@@ -0,0 +1,253 @@
+# Far-Edge Subcloud Installation and Remediation Guide
+
+## Installation
+
+A Wind River subcloud installation normally runs with no manual
+intervention. In the event that an installation needs to be started
+manually:
+
+- Visit https://middleware.vcpfe.vzwops.com/caas/wrfix/ in a browser.
+
+- Enter the iLO IP address into the first text field (even though it's
+ labeled "cluster name").
+
+- Type ```wr``` in the playbook key field.
+
+- Click the submit button.
+
+If an OK message appears, the job has been successfully sent to
+ansible-queue. The remaining steps require an SSH shell on
+```root@vcpfe-queue-2-prod-birmingham```:
+
+- View ```/var/log/lisp/ansible-queue.log``` to ensure that the
+ playbook started.
+
+- View the specific ansible log in ```/var/log/ansible/queue/``` to
+ watch the playbook's progress.
+
+If you want to watch the OS installation (the first stage of the
+overall deployment), attach to the iLO console:
+
+- Find the subcloud data (see below for the query) and find the
+ ```ilo_host_address``` and ```bmc_password```.
+
+- ```# ssh bladecenter-lb-vip.meter.vzwops.com```
+
+- ```# do-bmc <ilo_host_address>```
+
+- Enter the password when prompted.
+
+- Type ```vsp``` to start the serial console.
+
+## Remediation
+
+All too often it is necessary to remediate failed Wind River
+installations. Even successful ones need a visit by an engineer to
+ensure that they are in fact complete. It is not possible to discover
+this information from the middleware database, the subcloud, nor the
+central controller alone. All three of these data sources must be
+accessed. TODO: write a dashboard which aggregates the required
+information from the central controller, subcloud and database.
+
+Common causes of failure (this is not an exhaustive list):
+
+- OAM network unreachable.
+- MGMT network unreachable.
+- MGMT subnet overlaps with another already-installed subcloud.
+- The subcloud does not sync in time after coming online.
+- The platform-integ-apps application fails to apply.
+
+(This does not include servers that have issues which prevent them
+from advancing through the automation pipeline far enough to begin a
+Wind River installation. These include, but are not limited to,
+hardware problems and incorrect firmware versions.)
+
+### Inspecting a subcloud
+
+Three data sources were listed above. Let's visit each one in detail.
+
+#### Middleware database
+
+The middlware database is our record of what should be installed and
+whether or not we have marked it as complete. The ```caas_wrbatch```
+view and ```caas_wrinstallschedule``` table are important in this
+context.
+
+Here is a handy query to pull up all relevant information by cluster
+name:
+
+ select b.cluster_name || '.vcpfe.vzwops.com' as cluster_name,
+ b.*,
+ s.*
+ from caas_wrbatch b
+ left outer join caas_wrinstallschedule s
+ on b.cluster_id = s.cluster_id
+ inner join caas_location l
+ on b.fuze_spm_site_id = l.fuze_spm_site_id
+ where b.cluster_name in ('<cluster_name>')
+ order by b.parent_cluster_name asc,
+ b.cluster_name asc;
+
+(The expression in the first column is a convenience: it builds a FQDN
+for the subcloud that can be directly used in an SSH command from your
+laptop.)
+
+Here is a sample record showing a completed subcloud:
+
+ cluster_name | wsbomagj-d663366-001.vcpfe.vzwops.com
+ id | 423
+ cluster_id | 423
+ fuze_spm_site_id | 663366
+ parent_oam_vip_hostname |
+ parent_oam_vip_address | 2001:4888:a12:3221:106:290:0:10
+ parent_mgmt_address_range_start | 2001:4888:a12:3220:106:290:0:10
+ parent_mgmt_address_range_end | 2001:4888:a12:3220:106:290:0:ffff
+ parent_mgmt_default_gateway | 2001:4888:a12:3220:106:28::
+ parent_mgmt_subnet | 2001:4888:a12:3220::
+ pxe_mac_address | 48:df:37:e2:67:34
+ ilo_hostname | wsbomagj-663366-rh-pe0e910-001
+ ilo_host_address | 2001:4888:2a10:90fd:103:40a:0:e001
+ oam_hostname | wsbomagj-663366-rh-pe2e910-001
+ oam_host_address | 2001:4888:2a10:90fd:103:40a:0:400
+ oam_vip_address | 2001:4888:2a10:90fd:103:40a:0:f400
+ oam_default_gateway | 2001:4888:2a10:90fd:103:2a0::
+ mgmt_address_range_start | 2001:4888:2a10:982b:103:40a::
+ mgmt_address_range_end | 2001:4888:2a10:982b:103:40a:0:f
+ mgmt_default_gateway | 2001:4888:2a10:982b:103:2a0::
+ mgmt_subnet | 2001:4888:2a10:982b:103:40a::
+ host_vlan | 2845
+ oam_vlan | 410
+ mgmt_vlan | 420
+ cluster_name | wsbomagj-d663366-001
+ parent_cluster_name | wsbomagj-c319918-003
+ vendor_name | HPE
+ intel_nic_firmware_version | 1.2585.0
+ maint_window_p | f
+ namespace_id | 425
+ namespace_name | wsbomagj-663366vzwcvdu-y-ss-x-05690012222
+ bmc_username | k8sctl
+ bmc_password | <redacted>
+ id | 709
+ date_scheduled | 2020-11-19 19:15:16.885864+00
+ date_completed | 2020-12-04 19:28:17.644994+00
+ date_last_failed | 2020-12-03 20:47:18.199344+00
+ kirke_ticket_number | 709
+ kirke_ticket_status |
+ kirke_ticket_completed | 2020-11-20 02:28:30.908054+00
+ cluster_id | 423
+
+The key points:
+
+- ```pxe_mac_address``` is not null
+- ```intel_nic_firmware_version``` == 1.2585.0
+- ```date_completed``` is not null
+- ```kirke_ticket_number``` is not null
+
+The last point is the crucial one. Setting ```kirke_ticket_number```
+to something other than null is the way we manually mark the Wind
+River installation as complete. It must be done with a query similar
+to the following:
+
+ update caas_wrinstallschedule
+ set kirke_ticket_number = id
+ where id = 709;
+
+#### Central controller
+
+Use the following command to see the status of every subcloud attached
+to a central controller:
+
+ dcmanager subcloud list
+
+If the subcloud in question is complete and healthy from the central
+controller's point of view, it will look like the following. This does
+not necessarily mean that all is well. The subcloud itself must be
+visited and checked.
+
+ [sysadmin@controller-0 ~(keystone_admin)]$ dcmanager subcloud list | grep wsbomagj-d663366-001
+ | 546 | wsbomagj-d663366-001 | managed | online | complete | in-sync |
+
+#### Subcloud
+
+If the central controller says that the subcloud is offline, the first
+thing to check is the mgmt network. If you cannot ping the central
+controller's mgmt IP, there is no hope for remediating the subcloud.
+
+ ping6 <parent_mgmt_address_range_start>
+
+If the subcloud is managed, online and in-sync, then the remaining
+checks can be performed. Check the alarms:
+
+ fm alarm-list
+
+Check the applied status of the applications:
+
+ [sysadmin@controller-0 ~(keystone_admin)]$ system application-list
+ +--------------------------+----------+-----------------------------------+------------------------------------+---------+-----------+
+ | application | version | manifest name | manifest file | status | progress |
+ +--------------------------+----------+-----------------------------------+------------------------------------+---------+-----------+
+ | cert-manager | 20.06-4 | cert-manager-manifest | certmanager-manifest.yaml | applied | completed |
+ | nginx-ingress-controller | 20.06-0 | nginx-ingress-controller-manifest | nginx_ingress_controller_manifest. | applied | completed |
+ | | | | yaml | | |
+ | | | | | | |
+ | oidc-auth-apps | 20.06-26 | oidc-auth-manifest | manifest.yaml | applied | completed |
+ | platform-integ-apps | 20.06-9 | platform-integration-manifest | manifest.yaml | applied | completed |
+ +--------------------------+----------+-----------------------------------+------------------------------------+---------+-----------+
+
+Check the drbd resize. drbd8 should be set to 32GB:
+
+ [sysadmin@controller-0 ~(keystone_admin)]$ lsblk | grep drbd8
+ │ └─drbd8 147:8 0 32G 0 disk /var/lib/docker-distribution
+
+Check the PTP configuration. ```tx_timestamp_timeout``` should be set
+to 50:
+
+ [sysadmin@controller-0 ~(keystone_admin)]$ grep tx_timestamp_timeout /etc/ptp4l.conf
+ tx_timestamp_timeout 50
+
+Check the ISOL CPU configuration which should be enabled on cores 3,
+4, 27 and 28:
+
+ [sysadmin@controller-0 ~(keystone_admin)]$ system host-cpu-list controller-0 | grep -F 'Application-isolated'
+ | 72499a12-e2f2-40fc-a069-a2b002908843 | 3 | 0 | 3 | 0 | Intel(R) Xeon(R) Gold 6212U CPU @ 2.40GHz | Application-isolated |
+ | 8498669d-9048-45f6-b9dd-c81262934b70 | 4 | 0 | 4 | 0 | Intel(R) Xeon(R) Gold 6212U CPU @ 2.40GHz | Application-isolated |
+ | 4a70893f-b2cf-4405-9b69-11ed823a81b3 | 27 | 0 | 3 | 1 | Intel(R) Xeon(R) Gold 6212U CPU @ 2.40GHz | Application-isolated |
+ | 1d6e330d-7d2f-416c-b375-9b909d6a1f3c | 28 | 0 | 4 | 1 | Intel(R) Xeon(R) Gold 6212U CPU @ 2.40GHz | Application-isolated |
+
+### Marking an installation as complete
+
+If the subcloud passes all of the checks, it should be marked complete
+with the update query:
+
+ update caas_wrinstallschedule
+ set kirke_ticket_number = id
+ where id = <id>;
+
+### Remediating
+
+If the subcloud fails to pass all of the checks, it must be
+remediated. Deciding which steps to take requires operational
+experience. Here is a guide to some common problems and solutions:
+
+- The subcloud is unmanaged: ping the central controller mgmt IP. If
+ it fails, the network is misconfigured and must be fixed by the
+ region.
+
+- ```platform-integ-apps``` is not applied: perform a wipedisk and
+ reinstall.
+
+- ```oidc-auth-apps``` is uploaded but not applied: run the
+ ```wr_remediate``` playbook.
+
+- The drbd resize hasn't happened: run the ```wr_remediate```
+ playbook.
+
+- PTP isn't configured: run the ```wr_remediate``` playbook.
+
+- ISOL CPUS are not configured: run the ```wr_remediate``` playbook.
+
+The very last thing that the ```wr-installer``` playbook does is a
+```collect all```. You can tell if the playbook ran to completion by
+checking the contents of the ```/scratch``` directory on the
+subcloud. If it contains a collect tarball, the playbook completed.
+