Friday, March 20, 2015

NSX Host Preparation Not Ready

ShortURL to this blog post: http://vexpert.me/nsxhostpreptshoot

One of the steps required to install NSX-v is preparing clusters or hosts for NSX which will install three new VIBs to ESXi hosts (in the Kernel space) and the process will be done per cluster not per host.
The steps are quite straight forward, after you have your NSX Manager registered to your vCenter Server and get your NSX Controllers deployed, you just need to go to Host Preparation Tab > choose your desired cluster > hit Install


More about preparing clusters for NSX can be found from below links:
Getting Started Guide for NSX vSphere
NSX will send a request to EAM to calls vCenter Server to install the three VIBs.
ESXi downtime or reboot is not required for this but Upgrade or Uninstall requires a reboot, if you have DRS and vMotion networking setup, it will automatically set some of your hosts to maintenance and reboot it so there will be no downtime to your workload VMs.
Upon successful installation, the Installation Status column will display the NSX version and Firewall column Enabled.
But if not successful, the Installation Status column will display a red warning icon says Not Ready and Resolve.
Try clicking the resolve button to repush NSX request to install the VIBs, if nothing changed and still saying Not Ready you might need to check your setup and also check some logs.
Below are 2 log files that you can check:
  1. /var/log/esxupdate.log in ESXi host
  2. eam.log in vCenter Server. ProgramData/VMware/VMware VirtualCenter/Logs/ for Windows vCenter or /storage/log/vmware/vpx/eam.log for Linux vCenter
Below are the most common issues:
  1. DNS configuration on your vCenter Server, ESXi host, NSX Manager, and DNS forward + reverse records on your DNS servers
  2. Time Sync on ESXi Hosts and NSX Managers, use NTP server whenever possible
  3. TCP ports, check if any of required firewall ports are blocked. See the NSX Hardening Guide to check all required ports for NSX
  4. Verify if vCenter Server Managed IP is set and correct. See KB 1008030 on how to check this.
  5. Update Manager service is stopped or having issues if you have VUM installed. 
If you have error log under eam.log similar to below, you might have some issues on your Update Manager or connection to your Update Manager
“com.vmware.vim.vmomi.client.exception.ConnectionException: org.apache.http.conn.HttpHostConnectException: Connection to https://<vCenter_Server_Hostname>:8084 refused”
A workaround for this is to temporarily bypass VUM so EAM can continue to install the VIB. The steps are similar to KB 2053782 but instead of setting it to false, set it to true.

Below are the complete steps:
  1. Open a browser to the MOB at:
    https://vCenter_Server_IP/eam/mob/
  2. Log in with Administrator privileges.
  3. Select an agency from the agency list. This is usually agency-0. If not, select the appropriate agency.
  4. Click Config. This should be a URL where agency-0 is replaced by your local agency ID.
    For example:
    https://vCenter_Server_IP/eam/mob/?moid=agency-0&doPath=config
  5. Set the bypassVumEnabled flag to true:
  6. Using the agency ID identified in Step 3, change the browser URL:
    https://vCenter_Server_IP/eam/mob/?moid=agency-<0>&method=Update
  7. In the config field, change the value from false to true. Change the existing configuration options to match this:
    <config>
    <bypassVumEnabled>true</bypassVumEnabled>
    </config>

  8. Click Invoke Method and go back to the previous link to check if the value gets updated.
  9. Back to NSX Host Preparation Tab and click Resolve. You should be able to proceed with the VIBs installation now. Note: if you have multiple clusters, you would need to set the bypassVumEnabled flag to true per cluster since the agency-# will be created per cluster
Below are some related links that are helpful to understand NSX Host Preparation and help you further to troubleshoot:
 If you are still not able to resolve your issue, you can always contact VMware Support


Saturday, April 26, 2014

Solution - VMware vSphere Datastore is not accessible or not detected after unplanned PDL

I would like to share some KB & solutions related to PDL.
Some explanations on storage connectivity problems - PDL (Permanent Device Loss) and APD (All Path Down) is available in VMware Documentation here
Identifying Device Connectivity Problems
Permanent Device Loss (PDL) and All-Paths-Down (APD) in vSphere 5.x (2004684)

Some of my customers experiences PDL, unplanned PDL, where they can no longer access the datastore but storage team has confirmed that they have presented the LUN to the hosts & clusters.
However, the datastore is not detected or not accessible from vSphere.
Below are some of the scenario and its resolution.


1. After unplanned PDL, vSphere can see the Datastores. However, when we try to browse the datastore the Datastore Browser is just loading ... no response.
Solution: Follow this KB VMware KB: Cannot remount a datastore after an unplanned permanent device loss (PDL) or just simply reboot the ESXi host

2. After unplanned PDL, Datastore is gone from vSphere. Storage team confirmed that they have presented the LUN to the hosts & clusters. However, the datastore is not detected as existing datastore and also not detected as new LUN.
Solution: Check storage devices from the host, see if you found any detached devices

If you found any, check the naa id/LUN# if it is the LUN that you are looking for. Select the device then attach and rescan. The datastore you are looking for will normally appear in the list of datastore



3. After unplanned PDL, Datastore is gone but you can see it when you want to add new Datastore (detected as if it is a new Datastore).
You do not want to reformat the datastore but the options to keep or resignature VMFS only are greyed out. The available option is only to format VMFS.

Solution: Follow this KB VMware KB: Cannot remount a datastore after an unplanned permanent device loss (PDL)  or just simply reboot the ESXi host


4. After unplanned PDL,You want to remove inactive datastore because of PDL, but all options are greyed out
Delete it from vCenter Database, the steps are explained in this KB: VMware KB: Unable to remove a datastore from the vCenter Server 4.x / 5.x inventory
or you can also reboot the affected ESXi host

To avoid these issues follow the proper way/best practices for removing a LUN from ESX host
VMware KB: Unmounting a LUN or detaching a datastore/storage device from multiple VMware ESXi 5.x hosts
Best Practice: How to correctly remove a LUN from an ESX host | VMware vSphere Blog - VMware Blogs

Wednesday, April 16, 2014

VMware vCenter SQL Database Maintenance Best Practices

People often ask on what are best practices on maintaining vCenter SQL Database.
One of the VMware Best Practices (listed in VMware Health Check/HealthAnalyzer tool) is to periodically perform database maintenance tasks on the vCenter database.
When vCenter server service is stopped and cannot be started, most of the problem is because the disk which stores the DB data is full.
This is why monitoring the disk space and utilization is important to ensure that the database has sufficient space for growth.
We also should schedule regular backups of the vCenter database. The backup for vCenter Server should also include the SSL certificates and licenses from the vCenter Server.

vCenter stores configuration, tasks, events and performance data records in Database, the configuration record usually do not grow or changing most of the time, only happens when we change the setting of a cluster, adding host to cluster, etc. Tasks, events, and performance data records do grow over time and will populate table rows in Database as time goes by.

vCenter Server has a Database Retention Policy setting that allows you to specify when vCenter Server tasks and events should be deleted. vCenter has mechanism to purge the database so that it does not overgrow. There is some built-in vCenter SQL DB automated jobs in Microsoft SQL Server to clean performance data, tasks and events records in Database. Since the retention policy does not affect performance data records, it is still possible to purge or shrink old records from the database using the scripts available in this KB
Reducing the size of the vCenter Server database when the rollup scripts take a long time to run
http://kb.vmware.com/kb/1025914

To access the Database Retention Policy setting in the vSphere Client: Click Administration > vCenter Server Settings > Database Retention Policy.
If it's not set, then it means there is no imitation on how long vCenter will keep tasks and events records in the database, this can also lead to database overgrowth. The default setting is 180 days, so vCenter will purge old data after 180 days.

vCenter performs basic statistics operations of insert, roll up, and purge. Higher statistics levels require that more work be performed by the vCenter Server for these operations, which can impact the performance of the vCenter Server database.
Higher statistics levels also increase the size of the vCenter database. You can use the database sizing estimator when changing the statistics level to make sure that you have adequate space in the vCenter database.

vCenter statistics levels:
  1. vCenter statistics level 1 includes the basic metrics but does not include statistics for devices.
  2. vCenter statistics level 2 includes all the metrics including statistics for devices.
  3. vCenter statistics level 3 includes all the metrics and all of the counter groups.
  4. vCenter statistics level 4 includes all the metrics supported by vCenter Server.
To prevent performance data from growing so large, we can set the stats collection level to 1.
It is recommended that the vCenter statistics level is kept at level 1 or 2. Level 2 gives more comprehensive vCenter statistics than the default setting.
When increasing level statistics from default level 1, monitor closely the vCenter database growth.

Consider upgrading to vCenter 5.1+ if you are still using vCenter version prior v5.1.
vCenter Server 5.1 introduces some significant improvements to the statistics subsystem. The improvements are especially important for vCenter Server 5.1 deployments running at-scale inventory. The database is therefore a critical component of vCenter Server performance. Because the statistics data consumes a large fraction of the database, proper functioning of statistics is an important consideration for the overall database performance. Thus, statistics collection and processing are key components for vCenter Server performance.

VMware vCenter Server 5.1 Database Performance Improvements and Best Practices for Large-Scale Environments: https://www.vmware.com/files/pdf/techpaper/VMware-vCenter-DBPerfBestPractices.pdf

Below are some references & KBs related to vCenter Database Maintenance:

Friday, August 22, 2008

Recover IOS using tftpdnld from ROMMON

In Cisco 2600/2800/3800 Series Router we can recover IOS using Trivial File Transfer Protocol (TFTP) over ethernet interface using the ROMmon tftpdnld command.
tfptdnld is more faster rather than recovering IOS via Xmodem.

There are some variables to set when we want to transfer files to router using tftpdnld.
You can type tftpdnld -h

rommon 1 > tftpdnld -r

usage: tftpdnld [-hr]


Use this command for disaster recovery only to recover an image via TFTP.

Monitor variables are used to set up parameters for the transfer.

(Syntax: "VARIABLE_NAME=value" and use "set" to show current variables.)
"ctrl-c" or "break" stops the transfer before flash erase begins.


The following variables are REQUIRED to be set for tftpdnld:
IP_ADDRESS: The IP address for this unit
IP_SUBNET_MASK: The subnet mask for this unit
DEFAULT_GATEWAY: The default gateway for this unit

TFTP_SERVER: The IP address of the server to fetch from

TFTP_FILE: The filename to fetch


The following variables are OPTIONAL:
GE_PORT: Ethernet port number for download, 0 or 1 (default=0)

TFTP_MEDIA_TYPE: Media select for GE_PORT=0, 0(Copper) or 1(Fiber) (default=0)
TFTP_VERBOSE: Print setting. 0=quiet, 1=progress(default), 2=verbose
TFTP_RETRY_COUNT: Retry count for ARP and TFTP (default=20)
TFTP_TIMEOUT: Overall timeout of operation in seconds (default=7200)

TFTP_CHECKSUM: Perform checksum test on image, 0=no, 1=yes (default=1)

TFTP_MACADDR: The MAC address for this unit

GE_SPEED_MODE: 0=10/hdx, 1=10/fdx, 2=100/hdx, 3=100/fdx, 4=1000/fdx,
5=Auto (default)

Command line options:
-h: this help screen

-r: do not write flash, load to DRAM only and launch image


Below is an example in using tftpdnld to recover an IOS image name c3845-adventerprisek9-mz.124-21.bin to a Cisco 3845 router:

rommon 1 > IP_ADDRESS=171.68.171.0
rommon 2 > IP_ADDRESS=10.0.0.1

rommon 3 > IP_SUBNET_MASK=255.255.255.0

rommon 4 > DEFAULT_GATEWAY=10.0.0.13

rommon 5 > TFTP_SERVER=10.0.0.13

rommon 6 > TFTP_FILE=c3845-adventerprisek9-mz.124-21.bin

rommon 7 > tftpdnld

IP_ADDRESS: 10.0.0.1

IP_SUBNET_MASK: 255.255.255.0

DEFAULT_GATEWAY: 10.0.0.13

TFTP_SERVER: 10.0.0.13
TFTP_FILE: c3845-adventerprisek9-mz.124-21.bin
GE_PORT: Ge0/0
TFTP_MEDIA_TYPE: Copper

GE_SPEED_MODE: Auto
Invoke this command for disaster recovery only.

WARNING: all existing data in all partitions on flash will be lost!

Do you wish to continue? y/n: [n]: y


Receiving c3845-adventerprisek9-mz.124-21.bin from 10.0.0.13 !!!!!!!!!!!!!!!!!!!!!!!!!!!!

File reception completed.
Copying file c3845-adventerprisek9-mz.124-21.bin to flash.
Erasing flash at 0x607c0000

program flash location 0x60440000

rommon 8 >


References:
http://www.cisco.com/en/US/products/hw/routers/ps259/products_tech_note09186a008015bf9e.shtml