Showing posts with label RAC. Show all posts
Showing posts with label RAC. Show all posts

How to change Oracle 11gR2 RAC Database Files Layout and Public IP/VIP/SCAN IP etc

Jephe Wu - http://linuxtechres.blogspot.com

Objective:  Change all database files location for 11.2.0.3 database and RAC public ip/vip/scan ip etc
Environment: Oracle 11.2.0.3 RAC, 2 nodes, 2 VIPs, 3 SCAN IPs

References:
http://www.oracle-base.com/articles/misc/renaming-or-moving-oracle-files.php


Part I - Change Oracle 11gR2 RAC database files layout

1. check current database files info

sqlplus / as sysdba

sql> select name from v$datafile
sql> select name from v$tempfile;
sql> select member from v$logfile;
sql> archive log list;
sql> select dest_name, status, destination from v$archive_dest
sql> show parameter spfile;
sql> show parameter control;
sql> show parameter LOG_ARCHIVE_DEST_1;


sql> select * from database_properties


2. change all database files location - offline mode


---------control file and archive log file

sqlplus / as sysdba
sql> show parameter spfile;
sql> create pfile='/tmp/pfile' from spfile;
sql> shutdown immediate;

vi pfile to make changes for controlfiles and archive log location line
sqlplus / as sysdba
sql> startup nomount;
sql> create spfile='/u03/spfile/spfileDBID.ora' from pfile='/tmp/pfile';
sql> startup mount;
sql> alter database open;
sql> show parameter control_files;

---------online redo logs, data files, temp files

sqlplus / as sysdba
sql> shutdown immediate;
sql> startup mount;
sql> select name from v$datafile;
sql> select name from v$tempfile;
sql> select name from v$controlfile;
sql> select log_mode from v$database;
sql> select member from v$logfile;
sql> host
cp -va /u02/oradata/DBID/*.dbf /u03/oradata/DBID/
cp -va /u02/oradata/DBID/redo* /u04/oralog/DBID/
exit
sql> alter database rename file '/u02/oradata/DBID/file1.dbf' to '/u03/oradata/DBID/file1.dbf';
Note: do above for all output from 'select name from v$datafile;' and 'select name from v$tempfile;'

sql> alter database open;

3. change all database files location - online mode

------ control file
sql> show parameter control_files;

while database is running/online, run
sql>alter system set control_files = '/u09/control/DBNAME/control01.ctl', '/u09/control/DBNAME/control02.ctl','/u09/control/DBNAME/control03.ctl'

scope=spfile;

sql>shutdown immediate;

--------redo log

sqlplus / as sysdba
sql> select * from v$log;  # show thread and group numbers
sql> select a.group#,a.member,b.status,b.archived,bytes from v$logfile a, v$log b where a.group# = b.group# order by 1,2;
or
sql> select a.group#,a.member,b.status,b.archived,bytes/1024/1024 mbytes from v$logfile a, v$log b where a.group# = b.group# order by 1,2;
sql>

# check which thread for which node:
sql> select * from v$instance;
or
grep -i instance /tmp/pfile
or
strings spfilename | grep -i instance

add 2 more groups for each thread

ALTER DATABASE ADD LOGFILE [THREAD 1]
 GROUP 5 ('/u04/oralog/NRMAPS0/redo05.log') size 100M,
 GROUP 6 ('/u04/oralog/NRMAPS0/redo06.log') size 100M;


ALTER DATABASE ADD LOGFILE [THREAD 2]
 GROUP 7 ('/u04/oralog/NRMAPS0/redo07.log') size 100M,
 GROUP 8 ('/u04/oralog/NRMAPS0/redo08.log') size 100M;

alter database drop logfile group 1;

---------- data file
normal method:

If database is in archive log mode and the datafile is not system tablespace.
sql> ALTER DATABASE DATAFILE '/old/location' OFFLINE;
SQL> ALTER DATABASE RENAME FILE '/old/location' TO '/new/location';
SQL> RECOVER DATAFILE '/new/location';
SQL> ALTER DATABASE DATAFILE '/new/location' ONLINE;

You can online rename datafiles provided that datafile is not in SYSTEM tablespace.
sql> select name from v$tablespace;
sql> SELECT FILE_NAME, STATUS FROM DBA_DATA_FILES  WHERE TABLESPACE_NAME = 'USERS';
sql> alter tablespace users read only;
sql> SELECT TABLESPACE_NAME, STATUS FROM DBA_TABLESPACES where tablespace_name = 'USERS';
sql> host
use os utiliy to copy files from old location to new location
sql> alter tablespace users offline;
sql> alter database rename file 'old path' to 'new path';
sql> alter tablespace users online;
sql> alter tablespace users read write;
host
delete/move old files

rman method:

rman> report schema;
rman> copy datafile 3 to '/new/path/filename';
RMAN> SQL 'ALTER TABLESPACE xyz OFFLINE';

RMAN> SWITCH DATAFILE 3 TO COPY;
RMAN> RECOVER TABLESPACE xyz;


RMAN> SQL 'ALTER TABLESPACE soe ONLINE';
RMAN> HOST 'rm /old/path/filename';

rman> report schema;

-----------temp file

sql> SELECT v.file#, t.file_name, v.status from dba_temp_files t, v$tempfile v WHERE t.file_id = v.file#;
sql> ALTER DATABASE TEMPFILE '/path/to/file' offline;
sql> SELECT v.file#, t.file_name, v.status from dba_temp_files t, v$tempfile v WHERE t.file_id = v.file#;
sql> !cp -va /path/to/oldfile /path/to/newfile
sql> alter database rename file 'oldfile' to 'newfile';
sql> SELECT v.file#, t.file_name, v.status from dba_temp_files t, v$tempfile v WHERE t.file_id = v.file#;
sql> ALTER DATABASE TEMPFILE '/path/to/file' online;
sql> !rm -fr oldfile

------------archive log location

sql> set line 32000
sql> select dest_name, status, destination from v$archive_dest
sql> show parameter LOG_ARCHIVE_DEST_1;
sql> archive log list;
sql> alter system set log_archive_dest_1='LOCATION=/u05/oraarch/DBID';
sql> alter system archive log current;

4. change spfile, password file and ocr/votedisk etc

-------spfile

cd $ORACLE_HOME/dbs
normally, it's /u01/app/oracle/product/11.2.0/dbhome_1/dbs
vi initDBID.ora
put spfile location there like this:

SPFILE='/u12/spfile/spfileyourDBSID.ora'

then you need to change spfile location in clusterware as follows in oracle or root user:
/u01/app/11.2.0/grid/bin/srvctl modify database -d DBSID -p /u12/spfile/spfileyourDBSID.ora

------password file
create symbolic link under $ORACLE_HOME/dbs/orapwinstanceID pointing to centralized location for accessing from all nodes.

e.g.
database name is RACDB, two instance ID are RACDB1 and RACDB2

more $ORACLE_HOME/dbs/initRACDB1.ora
SPFILE='/u12/spfile/spfileRACDB.ora'

ls -l $ORACLE_HOME/dbs/orapwRACDB01.ora
orapwRACDB01.ora -> /u12/passwdfile/orapwRACDB

on another node, it's:

more $ORACLE_HOME/dbs/initRACDB2.ora
SPFILE='/u12/spfile/spfileRACDB.ora'

ls -l $ORACLE_HOME/dbs/orapwRACDB02.ora
orapwRACDB02.ora -> /u12/passwdfile/orapwRACDB

-------------ocr
online change ocr location:

touch ocrdisk;chown root:oinstall ocrdisk; chmod 640 ocrdisk  # make new ocr same permission as the existing ones

/u01/app/11.2.0/grid/bin/ocrcheck
/u01/app/11.2.0/grid/bin/ocrconfig -add /u12/crscfg/ocrdisk
/u01/app/11.2.0/grid/bin/ocrconfig -add /u12/crscfg/ocrdisk_mirror
/u01/app/11.2.0/grid/bin/ocrconfig -delete /u02/crscfg/ocr

------------votedisk
/u01/app/11.2.0/grid/bin/crsctl query css votedisk
crsctl query crs activeversion
/u01/app/11.2.0/grid/bin/crsctl add css votedisk /u12/crscfg/votedisk2
/u01/app/11.2.0/grid/bin/crsctl add css votedisk /u12/crscfg/votedisk3
/u01/app/11.2.0/grid/bin/crsctl delete css votedisk /u02/crscfg/vdsk

Part II - Change Public/Private/VIP/SCANIP etc

---------public ip
Refer to How to Modify Public Network Information Including VIP in Oracle Clusterware [ ID 276434.1 ]

You can change public NIC IP first then change ip in clusterware, or you can keep old ip, change clusterware info first, then change OS IP.

change OS IP first
then check and make sure clusterware is running by running commands below:
/u01/app/11.2.0/grid/bin/olsnodes -s
/u01/app/11.2.0/grid/bin/crsctl check clusterware -all

check public ip in OCR and clusterware:
in clusterware_interconnects parameter: oifcfg iflist -p -n
in OCR: oifcfg getif
from sqlplus : SELECT INST_ID, NAME_KSXPIA, IP_KSXPIA, PICKED_KSXPIA FROM X$KSXPIA;
debug interconnect traffic:
sqlplus / as sysdba
sql> oradebug setmypid
sql> oradebug ipc
then try to find the trace file under USER_DUMP_DEST parameter directory.
sql> show parameter user_dump_dest;
from sqlplus again:
sql> select * from v$cluster_initerconnects;
sql> select * from v$configured_interconnects;


oifcfg setif -global eth0/10.1.1.0:public
oifcfg delif -global eth0/10.2.1.0

--check network after changing public ip:
srvctl config network
srvctl modify network -k 1 -S 10.1.1.0/255.255.255.0/eth0

--stop/start clusterware:
srvctl stop cluster -all
srvctl start cluster -all
crs_stat -t
crsctl stat res -t

---------------VIP
Refer to: How to Modify Private Network Information in Oracle clusterware [ ID 283684.1 ]

check current VIPs
srvctl config nodeapps -a
crsctl stat res ora.db01.vip -p

use 'ifconfig eth0 192.168.2.1 netmask 255.255.255.0 up' to config OS ip first, then run
oifcfg setif -global eth3/192.168.2.0:cluster_interconnect

---------------SCANIP
refer to How to Modify SCAN settings or SCAN listener port after installation [ ID 972500.1 ]

crsctl stat res ora.scan1.vip -p
srvctl stop scan -f
crs_stat # to check scan name
srvctl modify scan -n db01-scan1
If SCANIP is using DNS name, you don't have to change it, just change DNS config.
if it's using /etc/hosts, and above command 'crsctl stat res ora.scan1.vip -p' shows you are using name for scan ip, you can add dummy line for

/etc/hosts for that scan ip

e.g.
root@db01:~# tail -3 /etc/hosts
#SCAN
10.12.1.200 db-scan.domain.com db-scan
10.12.1.200 db-scan1.domain.com db-scan1   # manually add this line

then modify it to db-scan1 first,then modify it back to db-scan
srvctl modify scan -n db-scan1
srvctl modify scan -n db-scan
then check it 'crsctl stat res ora.scan1.vip -p' , confirm it's using new IP address for SCAN IP

srvctl config scan
srvctl config scan_listener
srvctl modify scan_listener -u

Part III - Other RAC commands and troubleshooting

-------------misscount,reboottime and disktimeout 
refer to Steps To Change CSS Misscount, Reboottime and Disktimeout [ID 284752.1]
and
CSS Timeout Computation in Oracle Clusterware [ID 294430.1]

With 11gR2, these settings can be changed online without taking any node down:

1) Execute crsctl as root to modify the misscount:
     $CRS_HOME/bin/crsctl set css misscount <n>    #### where <n> is the maximum private network latency in seconds
     $CRS_HOME/bin/crsctl set css reboottime <r> [-force]  #### (<r> is seconds)
     $CRS_HOME/bin/crsctl set css disktimeout <d> [-force] #### (<d> is seconds)
2) Execute crsctl as root to confirm the change:
     $CRS_HOME/bin/crsctl get css misscount
     $CRS_HOME/bin/crsctl get css reboottime
     $CRS_HOME/bin/crsctl get css disktimeout

---------crs_stat -p
to check all detail settings for crs

crsctl enable crs  # enable startup for all crs daemons
crsctl disable crs
crsctl query crs softwareversion
crsctl query cdrs activeversion
crsctl check crs
crs_stat -t
crs_stat -p
crs_stat -ls
crsctl debug log res "ora.sdrac01.vip:5"
ocrconfig -showbackup

---------use srvctl to manage database resource
srvctl status database -d racdb
srvctl stop database -d racdb
srvctl start database -d racdb
srvctl stop instance -d racdb -i racdb1
srvctl start instance -d racdb -i racdb1

Note: it's recommended to use srvctl utility to manage database, otherwise, sometimes, if you use sqlplus to stop it, then its status is still online

in clusterware. you can use crs_stop to stop it then.

---------troubleshooting
select * from v$diag_info;  # show all alert and diagostic locations info


Part IV - find Oracle GI PSU and Oracle database latest patches


Bug 14727347 - 11.2.0.3.5 Grid Infrastructure Patch Set Update (GI PSU) [ID 14727347.8]
Patch 14727347: GRID INFRASTRUCTURE PATCH SET UPDATE 11.2.0.3.5 (INCLUDES DB PSU 11.2.0.3.5)



Preparing Oracle RAC system

Jephe Wu - http://linuxtechres.blogspot.com

Environment: openfiler as iscsi server, node1 and node2 are OL5.7, all of them are running under VirtualBox VM, Oracle 10gR2 10.2.0.1 clusterware

IP address assignment:
node1: eth0(public): 192.168.1.100, eth1(priv):172.16.1.100, vip:(going to be eth0:0) 192.168.1.102
node2: eth0(public): 192.168.1.101, eth1(priv):172.16.1.101, vip:(going to be eth0:0) 192.168.1.103
openfiler: 172.16.1.200 web login: https://172.16.1.200 (login as : openfiler/password)


Part I: OS installation
a. oracle user and groups
# groupadd oinstall
# groupadd dba
# groupadd oper
# useradd -u 200 -g oinstall -G dba[,oper] oracle

note: make sure all cluseter nodes has the same user and group id, otherwise ,it will fail to install clusterware.

b. verify nobody user exists
# id nobody

c. configure ssh on all nodes
make sure you can ssh into all nodes, public names and priv interconnect names, not vip ip address

d. check hardware requirements
memory size
swap size

e. NFS (http://download.oracle.com/docs/cd/B19306_01/install.102/b14203/prelinux.htm)
If you are using NFS for your shared storage, then you must set the values for the NFS buffer size parameters

rsize and wsize to at least 16384. Oracle recommends that you use the value 32768.

For example, if you decide to use rsize and wsize buffer settings with the value 16384, then update the /etc/fstab

file on each node with an entry similar to the following:

clusternode:/vol/DATA/oradata  /home/oracle/netapp     nfs    

rw,bg,vers=3,tcp,hard,nointr,timeo=600,rsize=32768,wsize=32768,actimeo=0  1 2


f. NTP for both nodes

g. /etc/hosts and dns
For each node, register one virtual host name and IP address in DNS.
For each private interface on every node, add a line similar to the following to the /etc/hosts file on all nodes,

specifying the private IP address and associated private host name:

h. /etc/sysctl.conf
kernel.shmall = 2097152

kernel.shmmax = 2147483648

kernel.shmmni = 4096

kernel.sem = 250 32000 100 128

fs.file-max = 65536

net.ipv4.ip_local_port_range = 1024 65000

net.core.rmem_default = 262144

net.core.rmem_max = 1048576

net.core.wmem_default = 262144

net.core.wmem_max = 1048576

run sysctl -p

i. Add the following lines to the /etc/security/limits.conf file:


oracle              soft    nproc   2047

oracle               hard    nproc   16384

oracle               soft    nofile  1024

oracle               hard    nofile  65536

j. Add or edit the following line in the /etc/pam.d/login file, if it does not already exist:


session    required     /lib/security/pam_limits.so
note: should not add above, according to redhat KB, otherwise you cannot login from console
Why the system text console cannot be login? - Article ID: 52661

the following should be added to /etc/profile:

if [ $USER = "oracle" ]; then

        if [ $SHELL = "/bin/ksh" ]; then

              ulimit -p 16384

              ulimit -n 65536

        else

              ulimit -u 16384 -n 65536

        fi

fi


k. identify oracle base  and home directories
more /etc/oraInst.loc  - Inventory directory
more /etc/oratab - Oracle home directory

l. create oracle base and clusterware home directory
mkdir -p /u01/app/oracle/
chown -R oracle:oinstall /u01/app/oracle
chmod -R 775 /u01/app/oracle 
(/u01 is the mount point)

mkdir -p /u01/crs/oracle/product/10/crs
chown root:oinstall /u01/crs/
chmod -R 775 /u01/crs/oracle


if mount point is /u01/, then the recommended Oracle clusterware home directory is
/u01/crs/oracle/product/10.2.1/crs


m. Verifying Hangcheck-timer Module on Kernel 2.6

For Red Hat Linux 4.0 and SUSE 9 systems, to verify that the hangcheck-timer module is running on every node:

    Enter the following command on each node to determine which kernel modules are loaded:

    # /sbin/lsmod

    If the hangcheck-timer module is not listed for any node, then enter a command similar to the following to

start the module located in the directories of the current kernel version:

    # insmod /lib/modules/kernel_version/kernel/drivers/char/hangcheck-timer.ko  hangcheck_tick=1

hangcheck_margin=10

    In the preceding command example, the variable kernel_version is the kernel version running on your system.

    To confirm that the hangcheck module is loaded, enter the following command:

    # lsmod | grep hang

    The output should be similar to the following:

    hangcheck_timer         3289  0

    To ensure that the module is loaded every time the system restarts, verify that the local system startup file

contains the command shown in the previous step, or add it if necessary:

        Red Hat:

        On Red Hat Enterprise Linux systems, add the command to the /etc/rc.d/rc.local file.

        SUSE:

        On SUSE systems, add the command to the /etc/init.d/boot.local file.


configure .bash_profile
export ORACLE_BASE=/u01/app/oracle
export ORACLE_HOME=$ORACLE_HOME/product/10.2.0/db_1
export ORA_CRS_HOME=$ORACLE_BASE/product/10.2.0/crs
export PATH=$PATH:$ORACLE_HOME/bin:$ORA_CRS_HOME/bin

2. clusterware installation

Please refer to http://oracleinstance.blogspot.com/2010/03/oracle-10g-installation-in-linux-5.html for complete

screenshot example
and also
http://space.itpub.net/21162451/viewspace-696413 - RHEL5.4+Oracle10gR2 RAC+OCFS2

2.1 os requirements

run as oracle for ./runInstaller to install 

Prerequisite Checks Fail When Installing 10.2 On Red Hat 5 (RHEL5)
Checking operating system version: must be redhat-3, SuSE-9, redhat-4, UnitedLinux-1.0, asianux-1 or asianux-2
                                      Failed <<<<

-> Prerequisite Checks Fail When Installing 10.2 On Red Hat 5 (RHEL5) [ID 456634.1]
If you are installing 10.2 from DVD, copy the <path>/database/install/oraparam.ini to a temporary directory (for

example, /tmp).

If you are installing 10.2 from an OTN download or have copied the 10.2 media to disk, take a backup of

<path>/database/install/oraparam.ini

Now edit oraparam.ini and change the appropriate line:

Original

[Certified Versions]
Linux=redhat-3,SuSE-9,redhat-4,UnitedLinux-1.0,asianux-1,asianux-2


New

[Certified Versions]
Linux=redhat-3,SuSE-9,redhat-4,UnitedLinux-1.0,asianux-1,asianux-2,redhat-5


note:
a. do not use ifconfig to configure vip before installing clusterware, just configure them in /etc/hosts
for example:

#cat /etc/hosts

127.0.0.1       localhost.localdomain localhost
192.168.1.100   jephe1.jephe.com jephe1
192.168.1.101   jephe2.jephe.com jephe2

172.16.1.100    jephe1-priv.jephe.com jephe1-priv
192.168.1.102   jephe1-vip.jephe.com jephe1-vip

172.16.1.101    jephe2-priv.jephe.com jephe2-priv
192.168.1.103   jephe2-vip.jephe.com jephe2-vip


OCR configration:
normal redundancy

specify OCR Location: /dev/raw/raw1
specify OCR Mirror Location: /dev/raw/raw2

When installing database software, you can choose 'configure ASM', then choose external for redundancy, then choose part of the raw disk for 'DATA' group, later we will run asmca to configure another disk group RECOVERY.

2.2 How To Setup UDEV Rules For RAC OCR And Voting Devices On SLES10, RHEL5, OEL5, OL5
- refer to http://www.held.org.il/blog/2007/11/setting-a-raw-device-in-redhatcentos-5/

configure /etc/sysconfig/rawdevices first, then service rawdevices restart to see those raw devices appear under /dev/raw/*

The following part are optional:
================
Add the required raw device ownership and permissions, for example:
a. Add to /etc/udev/rules.d/60-raw.rules:
ACTION==”add”, KERNEL==”sdb1″, RUN+=”/bin/raw /dev/raw/raw1 %N”
ACTION=="add", KERNEL=="sdc1", RUN+="/bin/raw /dev/raw/raw2 %N"
ACTION=="add", KERNEL=="sdd1", RUN+="/bin/raw /dev/raw/raw3 %N"
ACTION=="add", KERNEL=="sde1", RUN+="/bin/raw /dev/raw/raw4 %N"
ACTION=="add", KERNEL=="sdf1", RUN+="/bin/raw /dev/raw/raw5 %N"

==================
b. To set permission (optional, but required for Oracle RAC!), create a new /etc/udev/rules.d/99-raw-perms.rules

containing lines such as:

    KERNEL==”raw[1-5]“, MODE=”0640″, GROUP=”oinstall”, OWNER=”oracle”

Notice this:

    The raw-perms.rules file name has to begin with the number 99, which defines its order during rules apply, so that it will be used after all other rules take place. Using lower numbers might cause permissions to be incorrect.
    The following permissions have to apply:

OCR Device(s): root:oinstall , mode 0640
Voting device(s): oracle:oinstall, mode 0666
need to test the following commands:
# /sbin/udevcontrol reload_rules
# /sbin/start_udev


Or to add the following to 50-udev.rules
KERNEL=="vcsa[0-9]*",        NAME="%k", OWNER="vcsa", GROUP="tty", OPTIONS="last_rule"
KERNEL=="vcc/*",        NAME="%k", OWNER="vcsa", GROUP="tty", OPTIONS="last_rule"
KERNEL=="raw[1-9]",        OWNER="oracle", GROUP="oinstall", MODE="0640"
KERNEL=="raw10",        OWNER="oracle", GROUP="oinstall", MODE="0640"


# memory devices
KERNEL=="random",        MODE="0666", OPTIONS="last_rule"
KERNEL=="urandom",        MODE="0444", OPTIONS="last_rule"
KERNEL=="mem",            GROUP="kmem", MODE="0640", OPTIONS="last_rule"



permission is very important, otherwise, you might get error like this:
 when run 'crs_stat -t', node2 is offline and when run 'ps -efH' on node2, you will find 'startcheck'  like this:

root      99  416  0 1:23 ?        00:00:00 /bin/sh /etc/init.d/init.cssd startcheck
root     16  67  0 1:46 ?        00:00:00 /bin/sh /etc/init.d/init.cssd startcheck
root     1055  418  0 2:47 ?        00:00:00 /bin/sh /etc/init.d/init.cssd startcheck


check /var/log/messages, you will find this:
Feb  5 12:30:17 node2 logger: Cluster Ready Services waiting on dependencies. Diagnostics in /tmp/crsctl.3178.
Feb  5 12:31:17 node2 logger: Cluster Ready Services waiting on dependencies. Diagnostics in /tmp/crsctl.3353.
Feb  5 12:31:17 node2 logger: Cluster Ready Services waiting on dependencies. Diagnostics in /tmp/crsctl.3193.
Feb  5 12:31:17 node2 logger: Cluster Ready Services waiting on dependencies. Diagnostics in /tmp/crsctl.3178.
Feb  5 12:32:17 node2 logger: Cluster Ready Services waiting on dependencies. Diagnostics in /tmp/crsctl.3353.
Feb  5 12:32:18 node2 logger: Cluster Ready Services waiting on dependencies. Diagnostics in /tmp/crsctl.3193.


[root@node2 ~]# more /tmp/crsctl.3178
OCR initialization failed accessing OCR device: PROC-26: Error while accessing the physical storage Operating System error [Permission denied] [13]

check crs to confirm error:
/u01/oracle/product/crs/bin/crsctl check  crs
References:
http://surachartopun.com/2009/04/why-my-oracle-cluster-could-not-start.html

If one node is down, when running 'crs_stat -t', it will show this:

[oracle@node2 ~]$ crs_stat -t
Name           Type           Target    State     Host       
------------------------------------------------------------
ora....en1.gsd application    ONLINE    OFFLINE              
ora....en1.ons application    ONLINE    OFFLINE              
ora....en1.vip application    ONLINE    ONLINE    node2   
ora....en2.gsd application    ONLINE    ONLINE    node2   
ora....en2.ons application    ONLINE    ONLINE    node2   
ora....en2.vip application    ONLINE    ONLINE    node2 

Finally, if you encounter errors like this:

an error occurred during the interview for this component. oracle clusterware unable to retrieve voting disk information

then you need to 'rm -fr /u01/app/oracle/*' on node1 , then start ./runInstaller again.

2.3. at the end of clusterware installation, you will be prompted to run 2 programs which need root access permission.

permission,follow by the sequences below:
first program on first node
first program on second node

/home/oracle/oraInventory/orainstRoot.sh
ssh jephe2 -l root
/home/oracle/oraInventory/orainstRoot.sh
exit
/home/oracle/oracle/product/10.2.0/crs/root.sh
ssh jephe2 -l root
/home/oracle/oracle/product/10.2.0/crs/root.sh

second program on first node
second program on second node
basically, you need to finish the number 1 script on all nodes first before going to the next script.
also, run it under GUI interface, aka, X windows in vnc on node2 as root user, not oracle user, otherwise vipca will fail, vipca will use GUI

When you run the second /u01/app/oracle/product/10.2.0/crs/root.sh, there might be error like this:
 [root@oratest1 oratest1]# /u01/app/oracle/product/10.2.0/crs/root.sh
WARNING: directory '/u01/app/oracle/product/10.2.0' is not owned by root
WARNING: directory '/u01/app/oracle/product' is not owned by root
Checking to see if Oracle CRS stack is already configured

Setting the permissions on OCR backup directory
Setting up NS directories
Failed to upgrade Oracle Cluster Registry configuration
=>check log for detail: /u01/app/oracle/product/10.2.0/crs/log/oratest1/alertoratest1.log
see metalink Executing root.sh errors with "Failed To Upgrade Oracle Cluster Registry Configuration" [ID 466673.1]
 => solution is 
a) replace the biary according to above metalink
b) dd if=/dev/zero of=/dev/raw/raw1 bs=1024, (no need to finish, ctrl -c to cancel it)
run above root.sh again on node1


c. FAQ
c.1 some error at the end of running second root program on the second node:
Running vipca(silent) for configuring nodeapps
/home/oracle/oracle/product/10.2.0/crs/jdk/jre//bin/java: error while loading shared libraries: libpthread.so.0:

cannot open shared object file: No such file or directory

-> this is a bug, due to newer version of glibc(RHEL5) is incompatible with Java, need to modify vipca script as follows:

=> you can also modify it first before running root.sh on second node to avoid this error:

vi /home/oracle/oracle/product/10.2.0/crs/bin/vipca
change
LD_ASSUME_KERNEL=2.4.19
export LD_ASSUME_KERNEL  
     
to:              

LD_ASSUME_KERNEL=2.4.19
export LD_ASSUME_KERNEL                                          
unset LD_ASSUME_KERNEL


 #add this line to uncomment variable LD_ASSUME_KERNEL

if 'srvctl start XXX' command  reports same error, you can vi srvctl script to uncomment LD_ASSUME_KERNEL too, (which srvctl to find the path of the file)

Now, login xwindows on second node, run /home/oracle/oracle/product/10.2.0/crs/bin/vipca again to get errors below: (this error is unavoidable if you are using private ip range for public interface)

Error 0(Native: listNetInterfaces:[3])

[Error 0(Native: listNetInterfaces:[3])]


solution:

$ oifcfg iflist

eth0 192.168.1.0
eth1  172.16.1.0

$ oifcfg setif -global eth0/192.168.1.0:public
$ oifcfg setif -global eth1/172.16.1.0:cluster_interconnect
$ oifcfg getif


eth0 192.168.1.0  global  public
eth1 172.16.1.0  global  cluster_interconnect

# vipca  (run on the second node with root user login at X windows, don't run it with oracle user then su - as root)
just choose eth0 as vip interface, do not choose eth0:0, otherwise, after finishing installation, the vip will become eth0:0:1 for 192.168.1.102 for node1

You will be asked for entering ip alias name and ip address for both codegen1 and codegen2. Firstly, enter vip address for codegen1: 192.168.1.102, then the rest will come out automatically.
node1-vip.jephe.com and node2-vip.jephe.com are ip alias name.

after finishing configure vipca, back to CRS installation GUI, click ok to finish installation of clusterware

If environment variable
$ORA_CRS_HOME/cfgtoollogs/configToolFailedCommands.sh

If you got error: PRKR-1062 : Failed to find configuration for node codegen1,
then your /etc/hosts might be missing domain name such as
192.168.1.100 jephe1
not
192.168.1.100 jephe1.jephe.com jephe1

==============
cronjob for checking RAC configuration:

---------

#!/bin/sh
. /home/oracle/.bash_profile

DATE=`date +%Y%m%d`
TMPFILE=`mktemp`

echo "running ocrconfig -showbackup" > $TMPFILE
ocrconfig -showbackup >> $TMPFILE 2>&1

echo "" >> $TMPFILE
echo "running ocrcheck" >> $TMPFILE
ocrcheck >> $TMPFILE 2>&1

echo "" >> $TMPFILE
echo "rnning olsnodes -n -p -i" >> $TMPFILE
olsnodes -n -p -i >> $TMPFILE 2>&1

echo "" >> $TMPFILE
echo "running crsctl query css votedisk" >> $TMPFILE
crsctl query css votedisk >> $TMPFILE 2>&1

echo "" >> $TMPFILE
echo "running crsctl check crs" >> $TMPFILE
crsctl check crs >> $TMPFILE 2>&1

echo "" >> $TMPFILE
echo "running oifcfg iflist -p -n" >> $TMPFILE
oifcfg iflist -p -n >> $TMPFILE 2>&1

echo "" >> $TMPFILE
echo "running oifcfg getif" >> $TMPFILE
oifcfg getif >> $TMPFILE 2>&1


echo "" >> $TMPFILE
echo "running srvctl config nodeapps -n oradb-01 -a -g -s -l" >> $TMPFILE
srvctl config nodeapps -n oradb-01 -a -g -s -l >> $TMPFILE 2>&1

echo "" >> $TMPFILE
echo "running srvctl config nodeapps -n oradb-02 -a -g -s -l" >> $TMPFILE
srvctl config nodeapps -n oradb-02 -a -g -s -l >> $TMPFILE 2>&1



echo "" >> $TMPFILE
echo "cat /etc/oracle/ocr.loc" >> $TMPFILE
cat /etc/oracle/ocr.loc >> $TMPFILE

echo "" >> $TMPFILE
echo "crs_stat -t -v" >> $TMPFILE
crs_stat -t -v >> $TMPFILE


tar cpzf /tmp/crs.tar.gz /u01/crs/oracle/product/10.2.0/crs/cdata/crs

mutt -a /tmp/crs.tar.gz -s "crs/ocr status and backup on $DATE"  jwu@domain.com < $TMPFILE

rm -f $TMPFILE
---------------------


Part II: openfiler iscsi:

How to Dynamically Add and Remove SCSI Devices on Linux [ID 603868.1]


iscsiadm -m discovery -t sendtargets -p 192.168.1.5
 chkconfig iscsid on
chkconfig iscsi on

service iscsi resart
cd /var/lib/iscsi;ls


Reference: http://www.cyberciti.biz/tips/rhel-centos-fedora-linux-iscsi-howto.html



yum install lsscsi

lsscsi
cat /proc/scsi/scsi
grep host /etc/modprobe.conf
ls -ld /sys/class/scsi_host/host*
dmsetup ls | sort
multipath -d -ll
raw -qa
ls -l /dev/mapper/
ls -l /dev/raw/
ocrcheck
crsctl query css votedisk
crsctl check crs
ocrconfig -showbackup
cat /etc/oracle/ocr.loc

where is the setting for RAC service preference nodes?

nodes applications contain listener, gsnd etc?


Firstly, Oracle checks /etc/oracle/ocr.loc to find out where is the ocr (raw1 and raw2), then from ocr, to find out where are the voting disk devices.

[root@codegen1 bin]# lsscsi
[0:0:0:0]    disk    ATA      VBOX HARDDISK    1.0   /dev/sda
[2:0:0:0]    cd/dvd  VBOX     CD-ROM           1.0   /dev/sr0
[3:0:0:0]    disk    OPNFILER VIRTUAL-DISK     0     /dev/sdb
[3:0:0:1]    disk    OPNFILER VIRTUAL-DISK     0     /dev/sdc
[3:0:0:2]    disk    OPNFILER VIRTUAL-DISK     0     /dev/sdd
[3:0:0:3]    disk    OPNFILER VIRTUAL-DISK     0     /dev/sde
[3:0:0:4]    disk    OPNFILER VIRTUAL-DISK     0     /dev/sdf
[3:0:0:5]    disk    OPNFILER VIRTUAL-DISK     0     /dev/sdg
[3:0:0:6]    disk    OPNFILER VIRTUAL-DISK     0     /dev/sdh
[3:0:0:7]    disk    OPNFILER VIRTUAL-DISK     0     /dev/sdi
[3:0:0:8]    disk    OPNFILER VIRTUAL-DISK     0     /dev/sdj
[3:0:0:9]    disk    OPNFILER VIRTUAL-DISK     0     /dev/sdk

5.1. Resizing an Online Multipath Device
If you need to resize an online multipath device, use the following procedure.

    Resize your physical device.
    Use the following command to find the paths to the LUN:

    # multipath -l
[ -ll] [-d -ll]

    Resize your paths. For SCSI devices, writing a 1 to the rescan file for the device causes the SCSI driver to rescan, as in the following command:

    # echo 1 > /sys/block/device_name/device/rescan

    Resize your multipath device by running the multipathd resize command:

    # multipathd -k'resize map mpath0'

    Resize the filesystem (assuming no LVM or DOS partitions are used):

    # resize2fs /dev/mapper/mpath0

For further information on resizing an online LUN, see the Online Storage Reconfiguration Guide.

echo "scsi add-single-device 3 0 0 0" > /proc/scsi/scsi


Reference:
1. http://space.itpub.net/21162451/viewspace-696413 - RHEL5.4+Oracle10gR2 RAC+OCFS2
2. http://www.oracle.com/webfolder/technetwork/tutorials/demos/db/10g/r2/rac_r2_work/02_01_crs_install/02_01_crs_inst

all_viewlet_swf.html - Install 10g clusterware livedemo from Oracle
3. http://www.oracle.com/webfolder/technetwork/tutorials/demos/db/11g/r1/clusterware/installation_of_oracle_clusterware/installation_of_oracle_clusterware_viewlet_swf.html - Install clusterware on Oracle 11gR1
4. install oracle 11gR2 database - http://st-curriculum.oracle.com/obe/db/11g/r2/2day_dba/install/install.htm
5. http://oracleinstance.blogspot.com/2010/03/oracle-10g-installation-in-linux-5.html

Oracle RAC Documentation

Jephe Wu - http://linuxtechres.blogspot.com

  1. How to check if it's running RAC
BEGIN
  IF dbms_utility.is_cluster_database THEN
      dbms_output.put_line('Running in SHARED/RAC mode.');
  ELSE
      dbms_output.put_line('Running in EXCLUSIVE mode.');
  END IF;
END;
/
 
or
SQL> show parameter CLUSTER_DATABASE
 
 
 2. How to stop/start dbconsole?
Linux: 
$ export ORACLE_HOME=xxx
$ export ORACLE_SID=xxx
$ ORACLE_HOME/bin/emctl stop dbconsole
Windows:
set ORACLE_HOME=xxx
set ORACLE_SID=xxx
ORACLE_HOME\bin\emctl stop dbconsole 
 
Linux:
$ export ORACLE_HOME=xxx
$ export ORACLE_SID=xxx
$ ORACLE_HOME/bin/emctl start dbconsole
Windows:
set ORACLE_HOME=xxx
set ORACLE_SID=xxx
ORACLE_HOME\bin\emctl start dbconsole
Reference: Enterprise Manager Database Console FAQ [ID 863631.1]
 
3.  How to stop/start RAC 
 Bring up the inst1 of database db1 
$ srvctl start instance -d db1 -i inst1

 Stop the db1 database: all its instances and all its services, on all nodes.
$ srvctl stop database -d db1
 
 
4. database uptime
SELECT to_char(startup_time,'DD-MON-YYYY HH24:MI:SS') "DB Startup Time" FROM sys.v_$instance;
 
5. alert log location
show parameter BACKGROUND_DUMP_DEST; 
 
6. stop/start RAC
Starting the Oracle 10g RAC Cluster 10g Environment:

- Run as oracle:  su - oracle

$ export ORACLE_SID=orcl1
$ srvctl start nodeapps -n linux1

$ srvctl start asm -n linux1

$ srvctl start instance -d orcl -i orcl1

$ emctl start dbconsole

Start/Stop All Instances with SRVCTL



Stopping the Oracle 10g RAC Cluster 10g Environment: 

- Run as oracle:  su - oracle

$ export ORACLE_SID=orcl1

$ emctl stop dbconsole

$ srvctl stop instance -d orcl -i orcl1
srvctl status instance -d dbname -i instancename

$ srvctl stop asm -n linux1 [-o immediate]

$ srvctl stop nodeapps -n linux1

Starting the Oracle RAC 10g Environment

To stop or start both database instances at once: 

$ srvctl start database -d orcl [-o open | -o mount | -o nomount]

$ srvctl stop database -d orcl [-o normal | -o transactional | -o immediate | -o abort] 
srvctl status database -d dbname

srvctl config database -d dbname (shows instances name, node and oracle home) ==========
 
Services:
srvctl status service -d dbname

 srvctl config service -d dbname

 srvctl start service -d dbname -s servicename

srvctl stop service -d dbname -s servicename 
 
ï¼—. lsnrctl PLSExtProc, XDB and XPT
PLSExtProc: This is used to Call OS Commands from PL/SQL using External Procedures. Default in the Listener file. 
the XDB service is used for the XML DB database option 
the XPT service is used for Dataguard. 

All three of these can be disabled if you are not using their associated features.