Showing posts with label netapp. Show all posts
Showing posts with label netapp. Show all posts

Windows VM iscsi timeout when doing controller giveback for Netapp storage

Jephe Wu - http://linuxtechres.blogspot.com

Problem: Windows VM iscsi drive timeout and disappear after Netapp giveback operation, which caused Microsoft SQL server 2012 down, as data files are on iscsi D drive.
Objective:  find out why iscsi drive disappeared during giveback operation
Environment:  Windows 2008 R2 VM sitting on ESXi 5.1 cluster, D drive is iscsi drive on Netapp ONTAP 7 mode storage, VM is using Microsoft software iscsi initiator connecting Netapp portal target group.



Observation
When doing giveback between two controller heads on Netapp, iscsi drive on ESXi Windows 2008 R2 VM lost connection to the target which caused D drive timeout then disappear, which caused SQL server 2012 down. After around 18 minutes, iscsi drive reconnected back, SQL server restarted and operation resumed.

There was no any issues during takeover for Netapp, only there's issue during giveback.

Root Cause
After some research, the KB article from Netapp below indicated the problem:

Microsoft iSCSI SW Initiator takes a long time to reconnect to the filer after disruption. - http://support.netapp.com/NOW/cgi-bin/bol?Type=Detail&Display=202007

I've pasted above KB article as follows:
------------------

Bug ID202007
TitleMicrosoft iSCSI SW Initiator takes a long time to reconnect to the filer after disruption.
Duplicate of
Bug Severity3 - Serious inconvenience
Bug StatusClosed
ProductData ONTAP
Bug TypeISCSI - Windows 
Description Formatted
 During initial target discovery ("Add Target"), the Microsoft iSCSI SW
 Initiator uses the iSCSI SendTargets command, to retrieve from the target
 a list of IP addresses at which the target can be accessed.
 
 In a filer configuration with multiple physical networks or multiple VLANs,
 it is possible that some of the filer's addresses are not accessible to
 a given host.  In this situation, the SendTargets response sent by the
 filer will advertise some addresses which are not accessible by that host.
 
 When the Microsoft SW initiator loses connectivity to the target (such
 as during filer reboot, takeover, and giveback), the initiator attempts
 to restablish connectivity to the target using the following default
 algorithm, which Microsoft calls 'port-hopping':
 
   - attempt to reconnect over the same IP address which was being
     used before the disruption
   - cycle through the other IP addresses from the SendTargets response,
     attempting to reconnect, until connectivity is reestablished.
 
 Each inaccessible IP address in the list can add a delay of 15-20
 seconds, the TCP connection establishment timeout.  If there are many
 inaccessible IP addresses in the list, it may take a long time for the
 Microsoft initiator to cycle through the list before it finally
 successfully reconnect to the target.  If the total reconnect time
 exceeds the timeout configured on the host (MaxRequestHoldTime (non-MPIO),
 or PDORemovePeriod (MPIO)), the application will result in I/O errors.
 
Workaround Formatted
 This long reconnect time can be minimized by disabling the use of
 the Microsoft 'port-hopping' technique.  This is achieved by directing
 the Microsoft iSCSI initiator to use a specific IP address to create
 a TCP connection. The steps are:
 
 1. In the "Logon to target" box, click "Advanced ..."
 2. In the "Target Portal" list, change the value from "Default" to a
    specific IP address.  Select the same filer IP address as was
    originally specified in the "Add Target Portal" dialog under 
    "Discovery" tab.
 
 After the disruption occurs, the Microsoft initiator will use only the
 IP address previously specified in the advanced logon setting to reconnect
 to the filer and will not try other IP addresses advertised in the
 SendTargets response.
 
Notes Formatted
 
Fixed-In VersionThis bug is not scheduled to be fixed, you may opt to open a technical support case if you would like to contact NetApp regarding the status of this bug. A complete list of releases where this bug is fixed is available here.
Related Bugs
Bug Watch StatusThis bug is unwatchable.
--------------------

And also the article below explained above issue
http://software.tectrade.co.uk/SAN/NSeries/gc52129616.pdf  page 16 and 17

-------------
Microsoft iSCSI SW Initiator takes a long time to reconnect to the storage
system after disruption
In a storage system configuration with multiple physical networks or multiple
VLANs, the Microsoft iSCSI software Initiator can take several minutes to
reconnect to the storage system.

During initial target discovery (Add Target), the Microsoft iSCSI software
initiator uses the iSCSI SendTargets command to retrieve a list of IP addresses
at which the target can be accessed.

In a storage system configuration with multiple physical networks or multiple
VLANs, it is possible that some of the storage system's addresses are not
accessible to a given host. In this situation, the SendTargets response sent by
the storage system will advertise some addresses which are not accessible by
that host.

When the Microsoft iSCSI initiator loses connectivity to the target (such as
during storage system reboot, takeover, and giveback), the initiator attempts
to reestablish connectivity to the target using the following default algorithm,
which Microsoft calls “port-hopping”:

1 Attempt to reconnect over the same IP address which was being used
before the disruption.
2 Cycle through the other IP addresses from the SendTargets response,
attempting to reconnect, until connectivity is reestablished.
Each inaccessible IP address in the list can add a delay of 15-20 seconds,
which is the TCP connection establishment timeout. If there are many
inaccessible IP addresses in the list, it may take a long time for the iSCSI
initiator to cycle through the list before successfully reconnecting to the target.
If the total reconnect time exceeds the timeout configured on the host
(MaxRequestHoldTime for non-MPIO, or PDORemovePeriod for MPIO), the
Windows applications experience I/O errors.

This long reconnect time can be minimized by disabling the use of the
Microsoft “port-hopping” algorithm by using a specific target IP address for
each connection in the iSCSI Initiator.

Disabling the Microsoft iSCSI port-hopping algorithm
Disable the Microsoft iSCSI port-hopping algorithm in the iSCSI Initiator to
minimize the recovery time in configurations with many iSCSI target ports.
1. Open the Microsoft iSCSI initiator applet and select the Targets tab.
2. Click Log On.
3. In the Logon to target dialog box, click Advanced.
4. In the Target Portal list, change the value from Default to one of the
storage system IP addresses specified in the Target Portals list on the
Discovery tab.

After a disruption occurs, the Microsoft initiator uses only the IP address
specified to reconnect to the storage system and does not try other IP
addresses advertised in the SendTargets response.

Note: If the specified IP address is unreachable, no failover occurs.

---------------------

Note: Changing from 'default' to specific IP in advanced setting will be persistent across the reboot, see Microsoft support replied as follows:

----------
Once you have set the ISCSI configuration to use the specific IP address. Even though if we restart the server or the session gets disconnected the bindings are persistent. Unfortunately there is no specific utility by which we can confirm that it’s still using the Specific IP, but as per the configuration it will remain as it is.
This setting will remain persistent until and unless the entries from the Discovery Tab is removed. Once the entry from Discovery Tab is removed. We will have to re-configure once again with the specific IP`s.
------------------


How to Prove above solution from Netapp KB 
Install Wireshark on Windows VM, enable all interfaces monitoring and filter traffic only for iscsi (tcp.port==3260) During giveback operation. You will notice it will try all the IPs from non-stroage facing Interface if it cannot communicate with the correct IP in certain period
References

What are the parameters that control how MS iSCSI survives lost TCP connections without causing applications harm?

==========
During transient loss of connection, instead of reporting "Device not available" immediately, the Microsoft iSCSI initiator will try to reconnect to the target and resubmit outstanding SCSI commands.
There are three registry values related MS iSCSI retry behavior, found in the following path:
HKEY_LOCAL_MACHINE\SYSTEM\CurrentControlSet\Control\Class\{4D36E97B-E325-11CE-BFC1-08002BE10318}\[Instance_Number]\Parameters
Note: the [Instance_Number] may be different from system to system, depending on how many SCSI adapters already exist on the system.
    The three registry values are:
  1. DelayBetweenReconnect [default: 5 (seconds)]
  2. MaxConnectionRetries [default: 0xFFFFFFFF, infinite]
  3. MaxRequestHoldTime [default: 60 (seconds)]
Explanations: Normally you don't need to modify DelayBetweenReconnect and MaxConnectionRetries. The MaxRequestHoldTime is probably the only one that you may want to change. It defines how long Microsoft iSCSI initiator should hold and retry outstanding commands, before notifying upper layer of a Device Removal event. This event usually causes I/O failures to applications using the iSCSI disk. MaxRequestHoldTime is only relevant with non-MPIO environments. When MPIO is involved, this value is ignored.
A Device Removal event can be bad for applications actively using an iSCSI Logical Unit Number (LUN), especially if a cable-pull, filer reboot, filer cluster failover, etc., takes more than MaxRequestHoldTime of 60 seconds to recover. Unless you have special requirement that need the retry window to be smaller or larger, 180 (seconds) is a good value to start with.
Note: Even after a Device Removal event is reported, Microsoft iSCSI initiator will still keep trying to reconnect to the target, as defined by the first two registry values,DelayBetweenReconnect and MaxConnectionRetries.
The Windows iSCSI host must be rebooted after changing the registry value(s).
=============






How to make netapp Oracle database snapshot copy crash-consistent


Jephe Wu - http://linuxtechres.blogspot.com

Objective: understanding Netapp point-in-time snapshot Oracle backup without putting in the hot backup mode


Crash-consistent snapshot copies should only be considered under special circumstances where requirements restrict the use of standard backup methods (such as rman or hot backup mode)

Oracle backup overview:
----------------------
physical backup and logical backup

physical backup can be classified as consistent backup or inconsistent backup

consistent backup means controlfile and data file are checkpointed with same SCN, only possible when database is cleanly shut down, no matter it's in nonarchivelog or archivelog mode

Besides the standard 3 methods for backup: cold/offline backup, rman backup and online/hot backup(user-managed backup), Oracle recently certify the third party snapshot copy technology as one of options of backup/recovery as long as it's crash consistent

In the past, Oracle did not support or recommend the use of a snapshot copy created of an online active
database without the database or tablespaces being put in backup mode. The risk was thought to be
the danger of mixing old archive logs with current archive logs, which can lead to data corruption or
potentially destroy the production database.

According to MOS note 604683.1, the snapshot of an online database not in backup mode can be
deemed valid and supported if and only if all of the following requirements are strictly satisfied:
•  Oracle’s recommended restore and recovery operations are followed. 
•  Database is crash consistent at the point of the snapshot. 
•  Write ordering is preserved for each file within a snapshot. 


Oracle recovery overview:
------------------------
instance recovery and media recovery
instance recovery is automatic done by Oracle itself, it requires redo log file only.
Media recovery requires archived redo log.
Media recovery has complete recovery and incomplete recovery


An incomplete recovery of the whole database is usually required in the following situations:
•  Data loss caused by user errors
•  Missing archived redo log, which prevents complete recovery
•  Physical loss or corruption of online redo logs
•  No access to current control file

When performing incomplete recovery, the types of media recovery are available.


  • Time-based recovery  Recovers the data up to a specified point in time. 
  • Cancel-based recovery  Recovers until you issue the CANCEL statement (not available when using Recovery Manager). 
  • Change-based recovery  Recovers until the specified SCN. 
  • Log sequence recovery  Recovers until the specified log sequence number (only available when using Recovery Manager). 

What's the crash consistent?
------------------------------
It's point-in-time(PIT) image of Oracle database, looks like it crashed due to power outage, instance crash or shutdown abort etc, it requires instance recovery after restart database, not media recovery. Netapp snapshot generate Point-In-time image for database.


How to make snapshot crash-consistent?
---------------------------------------
1. all databqase files(controlfile, datafile, online redo log) are in single volume, then snapshot will generate crash-consistent image.
Note: Not require archived logs to be in the same volume.

If a database has all of its files (control files, data files, online redo logs, and archived logs) contained
within a single NetApp volume, then the task is straightforward. A Snapshot copy of that single volume
will provide a crash-consistent copy.

2. use crash consistent group by snapmanager/snapdrive etc if database cross different volumes
e.g. data volume and log volume, data captured by snapshot for data volume must exist in log volume first because Oracle always makes sure it writes to redo log first before writing associated data buffer cache to data file.


Starting from SnapDrive for unix 2.2, SnapDrive supports the feature of consistency groups provided
by Data ONTAP (beginning with version 7.2 and higher). This feature is necessary for creating a
consistent Snapshot copy across multiple controller/volumes.

In an environment where all participating controllers support consistency groups, SnapDrive will use a
Data ONTAP consistency group as the preferred (default) method to capture multicontroller/volume
Snapshot copies.

SnapDrive can simplify the creation of a consistency group Snapshot copy when there are
multiple file systems.

snapdrive snap create -fs /u01/oradata/prod /u02/oradata/prod -snapname snap_prod_cg 


a. POINT-IN-TIME COPY OF THE DATABASE 

After the database is opened, no future redo logs beyond this snapshot
can be applied.

Open resetlogs operation is recommended to avoid potential mixing of existing
archive logs and new archive logs. and start a new incarnation and log ID:

1. SHUTDOWN IMMEDIATE 
2.  STARTUP MOUNT 
3.  RECOVER DATABASE UNTIL CANCEL 
4.  ALTER DATABASE OPEN RESETLOGS; 

b. FULL DATABASE RECOVERY WITH ZERO DATA LOSS 

Restore the snapshot of only the data files. Do not overwrite the current control files, current redo
logs, and current archived logs.

run commands below to fully recover database by applying archived and online redo logs
1. recover automatic database;
2. alter database open;

c. point-in-time(PIT) database recovery
PIT requires the presence of current controlfile, current online redo logs and archived logs.

only restore data files and run the following commands:
1. startup mount

Identify the minimum SCN we have to recover to by script @scandatafile.sql

SQL> @scandatafile  
File 1 absolute fuzzy scn = 861391  
File 2 absolute fuzzy scn = 0  
File 3 absolute fuzzy scn = 0  
File 4 absolute fuzzy scn = 0  
Minimum PITR SCN = 861391  

PL/SQL procedure successfully completed. 


scandatafiles.sql 
# scans all files and update file headers with meta information  
# depending on number and sizes of files, the scandatafile procedure can be 
a  
# time consuming operation.  
# create a script, “scandatafile”, with the following content 

spool scandatafile.sql  
set serveroutput on  
declare  
 scn number(12) := 0;  
 scnmax number(12) := 0;  
begin  
 for f in (select * from v$datafile) loop  
 scn := dbms_backup_restore.scandatafile(f.file#);  
 dbms_output.put_line('File ' || f.file# ||' absolute fuzzy scn = ' || 
scn);  
 if scn > scnmax then scnmax := scn; end if;  
 end loop;  

 dbms_output.put_line('Minimum PITR SCN = ' || scnmax);  
end; 
/ 

If the minimum PITR SCN is zero, then database is not required for further recovery, it can to opened now.
if it's no zero, database must be recovered to at least that SCN and onwards.

2. RECOVER AUTOMATIC DATABASE UNTIL CHANGE [Minimum PITR SCN or higher]  
or 
ALTER DATABASE RECOVER DATABASE UNTIL CHANGE [Minimum PITR SCN or higher] 

3. ALTER DATABASE OPEN RESETLOGS 

References:

--------------
1. Using Crash-Consistent Snapshot Copies as Valid Oracle Backups - http://media.netapp.com/documents/tr-3858.pdf
2. MOS Supported Backup, Restore and Recovery Operations using Third Party Snapshot Technologies [ID 604683.1]

NFS client mount options for Oracle and Netapp

Jephe Wu - http://linuxtechres.blogspot.com

Environment: Linux x86_64, Oracle 11g 64bit SI(Single Instance) or RAC,  use Netapp as storage for Oracle binary, datafile, rman backup and expdp
Kernel 2.6, OL 5.8 and OL 6.3

Objective: the recommended NFS mounting options


Steps:

Please refer to Mount Options for Oracle files when used with NFS on NAS devices [ID 359515.1]

ORA-27054: NFS file system where the file is created or resides is not mounted with correct options [ID 781349.1]

and Netapp doc ID:
What are the mount options for databases on NetApp NFS?


Mount options for binary (exact order is important!)
RAC: rw,bg,hard,nointr,rsize=32768,wsize=32768,tcp,vers=3,timeo=600, actimeo=0
Netapp RAC: rw,bg,hard,rsize=32768,wsize=32768,vers=3,actimeo=0, nointr, suid, timeo=600, tcp

SI:    rw,bg,hard,rsize=32768,wsize=32768,vers=3,nointr,timeo=600,tcp
Netapp SI: rw,bg,hard,rsize=32768,wsize=32768,vers=3,nointr,timeo=600, tcp


For datafiles:
RAC: rw,bg,hard,nointr,rsize=32768, wsize=32768,tcp,actimeo=0, vers=3,timeo=600
Netapp RAC: rw,bg,hard,rsize=32768,wsize=32768,vers=3, actimeo=0, nointr, suid, timeo=600, tcp

SI:    rw,bg,hard,rsize=32768,wsize=32768,vers=3,nointr,timeo=600,tcp
Netapp SI: rw,bg,hard,rsize=32768,wsize=32768,vers=3,nointr, timeo=600, tcp

Mount options for CRS voting disk and OCR [RAC]
RAC: rw,bg,hard,nointr,rsize=32768, wsize=32768,tcp,noac,vers=3,timeo=600,actimeo=0
Netapp RAC: rw,bg,hard,rsize=32768,wsize=32768,vers=3,actimeo=0, nointr, suid, timeo=600, tcp

Note: 
For RMAN backup sets, image copies, and Data Pump dump files, the "NOAC" mount option should not be specified - that is because RMAN and Data Pump do not check this option and specifying this can adversely affect performance.

Due to Unpublished bug 5856342, it is necessary to use the following init.ora parameter when using NAS with all versions of RAC on Linux (x86 & X86-64 platforms) until 10.2.0.4. This bug is fixed and included in 10.2.0.4 patchset.
filesystemio_options = DIRECTIO

NOTE:  As per Bug 11812928, the 'intr' & 'nointr' are deprecated in UEK kernels, as well as Oracle Linux 6. It is harmless to still include it, you will get a notice..

NFS: ignoring mount option: nointr.



Variable Details: 

OptionDescription
hard Generate a hard mount of the NFS file system. If the connection to the server is lost temporarily, Oracle continues to retry the connection until the NAS device responds. 
bgTry to connect in the background if connection fails. 
proto=tcp(or tcp on Linux)Use the TCP protocol rather than UDP. TCP is more reliable than UDP. 
vers=3(or nfsvers=3 on Linux)Use NFS version 3. Oracle recommends that you use NFS version 3 where available, unless the performance of version 2 is higher. 
suid Allow clients to run executables with SUID enabled. This option is required for Oracle software mount points. 
rsize, wsize The number of bytes used when reading or writing to the NAS device. A value of 8192 is often recommended for NFS version 2 and 32768 is often recommended for NFS version 3. 
nointr (or intr)Do not allow (or allow) keyboard interrupts to kill a process that is hung while waiting for a response on a hard-mounted file system. 
noac
actimeo=0
Disable attribute caching. (a combination of sync and actimeo=0)
disable attribute caching on the client