Showing posts with label Exadata. Show all posts
Showing posts with label Exadata. Show all posts

Saturday, November 14, 2015

Patchmgr to upgrade Exadata database Nodes.


Updating Database Nodes with patchmgr
--------------------------------------
*) Starting with Exadata release 12.1.2.2.0, Oracle Exadata database nodes (running
releases later than 11.2.2.4.2), Oracle Exadata Virtual Server nodes (dom0), and Oracle
Exadata Virtual Machines (domU) can be updated, rolled back, and backed up using
patchmgr. You can still run dbnodeupdate.sh in standalone mode, but performing the
updates by running patchmgr enables you to run a single command to update
multiple nodes at the same time; you do not need to run dbnodeupdate.sh separately
on each node. patchmgr can update the nodes in a rolling or non-rolling fashion.

For patchmgr to do this orchestration, you run it from a database node that will not be updated itself.

*) Getting and Installing dbserver.patch.zip

Starting with release 12.1.2.2.0 a new dbserver.patch.zip file will be available for
running dbnodeupdate from patchmgr. The zip file contains patchmgr and
dbnodeupdate.zip. Unzip the dbserver.patch.zip file and run patchmgr from there.
Do not unzip dbnodeupdate.zip. Always check My Oracle Support note 1553103.1 for
the latest release of dbserver.patch.zip.

or Download .

Patch 21634633: DBSERVER.PATCH.ZIP ORCHESTRATOR PLUS DBNU - ARU PLACEHOLDER

*) When using the ISO file for the update, it is recommended that you put the ISO file in the same directory
where you have dbnodeupdate.zip.

Behavior for Non-Rolling Upgrades
The behavior for non-rolling upgrades is as follows:
- If a node fails at the pre-check stage, the whole process fails.
- If a node fails at the patch stage or reboot stage, patchmgr skips further steps for
the node. The upgrade process continues for the other nodes.
- The pre-check, patch/reboot, and complete stages are done in parallel.

The notification alert sequence is:
1. Started (All nodes in parallel)
2. Patching (All nodes in parallel)
3. Reboot (All nodes in parallel)
4. Complete Step (Each node serially)
5. Succeeded (Each node serially)

Updating database nodes using patchmgr is optional. You can still run
dbnodeupdate.sh manually. If you get any blocking errors from patchmgr when
updating critical systems, it is recommended that you perform the update by running
dbnodeupdate.sh manually.

./patchmgr -dbnode dbnode_list -dbnode_precheck -dbnode_loc patch_file_name -dbnode_version version
./patchmgr -dbnode dbnode_list -dbnode_precheck -dbnode_loc http://yum-repo/yum/ol6/EXADATA/dbserver/12.1.2.2.0/base/x86_64/ -dbnode_version 12.1.2.2.0.date_stamp
./patchmgr -dbnode dbnode_list -dbnode_precheck -dbnode_loc ./repo.zip -dbnode_version 12.1.2.2.0.date_stamp

./patchmgr -dbnode dbnode_list -dbnode_backup

./patchmgr -dbnode dbnode_list -dbnode_upgrade -dbnode_loc patch_file_name -dbnode_version version
./patchmgr -dbnode dbnode_list -dbnode_upgrade -dbnode_loc http://yum-repo/yum/ol6/EXADATA/dbserver/12.1.2.1.0/base/x86_64/ -dbnode_version 12.1.2.2.0.date_stamp -rolling
./patchmgr -dbnode dbnode_list -dbnode_upgrade -dbnode_loc ./repo.zip -dbnode_version 12.1.2.2.0.date_stamp -smtp_from "
Example of a rolling update using yum http repository:
# ./patchmgr -dbnode dbnode_list -dbnode_upgrade -dbnode_loc http://yum-repo/yum/ol6/EXADATA/dbserver/12.1.2.1.0/base/x86_64/ -dbnode_version 12.1.2.2.0.date_stamp -rolling

Example of a non-rolling update using zipped yum ISO repository:
# ./patchmgr -dbnode dbnode_list -dbnode_upgrade -dbnode_loc ./repo.zip -dbnode_version 12.1.2.2.0.date_stamp

Wednesday, July 15, 2015

Exadata patching Bios boot order is incorrect


in April - 2015 PSU , when we upgrade database node upgrade and again do ./dbnodeupdate.sh -c.

we may get below warning. to encounter it.

Warning: Bios boot order is incorrect - the system may have booting issues in a next reboot

if we encounter warning then do

ubiosconfig list status

if status pending then execute below command.

ubiosconfig cancel config

ubiosconfig list status

check status will be OK from Pending.

Monday, May 18, 2015

Exadata X5 Hybrid Rack | OVM and Physical Database Nodes in One Rack.

For my own curiosity, I was creating OEDA for latest Exadata X5 half rack and found that we can have Physical compute nodes and OVM Compute nodes in same cluster & in Same Rack.



Wednesday, May 13, 2015

Change Exadata Write-Back Flash cache in Rolling Mode.

---Make sure you have both of this files into one directory.

 [root@dummyCNadm01 ~]# pwd  
 /root  
 rwxr--r-- 1 root root 141286 May 12 13:47 setWBFC.sh  
 -rwxr-xr-x 1 root root  683 May 12 13:47 wbfc_FLUSH.sh  

----From compute Node. Run Pre-check to see if cell servers are ready to flip over WBFC .

 [root@dummyCNadm01 ~]# ./setWBFC.sh -g cell_group -l /tmp -m WriteBack -o rolling -p  
 ./setWBFC.sh: Using log directory '/tmp'  
 ./setWBFC.sh: Log File '/tmp/setWBFC_85166_2015-05-12-13:50:03.log' created successfully  
 2015-05-12 13:50:03  
 Starting ./setWBFC.sh on dummyCNadm01  
 Version: 1.0.0.1.6.20140716  
 Command line options used:  
  -g cell_group  
  -o rolling  
  -m WriteBack  
  -p (Perform pre-req checks only)  
  -t 21600  
  -x 0  
 2015-05-12 13:50:03  
 Performing pre-req checks.....  
 2015-05-12 13:50:03  
 Creating baseline inventory for griddisks  
 2015-05-12 13:50:06  
 Creating baseline inventory for flashdisks  
 2015-05-12 13:50:08  
 Creating baseline inventory for flashsize  
 2015-05-12 13:50:10  
 dcli present and in PATH.            [PASSED]  
 2015-05-12 13:50:10  
 Checking cell nodes are valid storage servers...  
 2015-05-12 13:50:10  
 All cells are valid Exadata storage cells.  
 2015-05-12 13:50:10  
 Checking Exadata Storage Software Versions...  
 2015-05-12 13:50:15  
 Software versions of the following cells:  
 dummy01celadm01: 12.1.2.1.1.150316.2           [PASSED]  
 dummy01celadm02: 12.1.2.1.1.150316.2           [PASSED]  
 dummy01celadm03: 12.1.2.1.1.150316.2           [PASSED]  
 dummy01celadm04: 12.1.2.1.1.150316.2           [PASSED]  
 dummy01celadm05: 12.1.2.1.1.150316.2           [PASSED]  
 dummy01celadm06: 12.1.2.1.1.150316.2           [PASSED]  
 dummy01celadm07: 12.1.2.1.1.150316.2           [PASSED]  
 2015-05-12 13:50:15  
 Checking Grid Infrastructure Software Version...  
 2015-05-12 13:50:17  
 Grid Infrastructure version: 12.1.0.2.00     [PASSED]  
 2015-05-12 13:50:17  
 Checking for active ASM operations....  
 2015-05-12 13:50:18  
 Check for no active ASM operations:       [PASSED]  
 2015-05-12 13:50:18  
 Checking griddisk status across all cells....  
 2015-05-12 13:50:20  
 All griddisks across all cells have asmdeactivationoutcome = Yes  
 All griddisks across all cells are ONLINE  
 Griddisk checks:                 [PASSED]  
 2015-05-12 13:50:20  
 Checking flash cache status.....  
 2015-05-12 13:50:21  
 Flashcache status normal:            [PASSED]  
 2015-05-12 13:50:21  
 Checking that all FlashDisks are present...  
 2015-05-12 13:50:22  
 FlashDisk validation:              [PASSED]  
 2015-05-12 13:50:22  
 Checking current flash cache mode.....  
 2015-05-12 13:50:23  
 Flashcache not already in target mode:      [PASSED]  
 2015-05-12 13:50:23  
 All pre-req checks completed:          [PASSED]  
 2015-05-12 13:50:24  
 dummy01celadm01: flashcache size: 5.82122802734375T  
 dummy01celadm02: flashcache size: 5.82122802734375T  
 dummy01celadm03: flashcache size: 5.82122802734375T  
 dummy01celadm04: flashcache size: 5.82122802734375T  
 dummy01celadm05: flashcache size: 5.82122802734375T  
 dummy01celadm06: flashcache size: 5.82122802734375T  
 dummy01celadm07: flashcache size: 5.82122802734375T  
 There are 7 storage cells to process.  

---- Turn write back ON in rolling fashion .

 [root@dummyCNadm01 ~]# ./setWBFC.sh -g cell_group -l /tmp -m WriteBack -o rolling  
 2015-05-12 14:59:41  
 Starting ./setWBFC.sh on dummy01dbadm01  
 Version: 1.0.0.1.6.20140716  
 Command line options used:  
  -g cell_group  
  -o rolling  
  -m WriteBack  
  -t 21600  
  -x 0  
 2015-05-12 14:59:41  
 Performing pre-req checks.....  
 2015-05-12 14:59:41  
 Creating baseline inventory for griddisks  
 2015-05-12 14:59:42  
 Creating baseline inventory for flashdisks  
 2015-05-12 14:59:43  
 Creating baseline inventory for flashsize  
 2015-05-12 14:59:44  
 dcli present and in PATH.            [PASSED]  
 2015-05-12 14:59:44  
 Checking cell nodes are valid storage servers...  
 2015-05-12 14:59:44  
 All cells are valid Exadata storage cells.  
 2015-05-12 14:59:44  
 Checking Exadata Storage Software Versions...  
 2015-05-12 14:59:48  
 Software versions of the following cells:  
 dummy01celadm01: 12.1.2.1.1.150316.2           [PASSED]  
 dummy01celadm02: 12.1.2.1.1.150316.2           [PASSED]  
 dummy01celadm03: 12.1.2.1.1.150316.2           [PASSED]  
 dummy01celadm04: 12.1.2.1.1.150316.2           [PASSED]  
 dummy01celadm05: 12.1.2.1.1.150316.2           [PASSED]  
 dummy01celadm06: 12.1.2.1.1.150316.2           [PASSED]  
 dummy01celadm07: 12.1.2.1.1.150316.2           [PASSED]  
 2015-05-12 14:59:48  
 Checking Grid Infrastructure Software Version...  
 2015-05-12 14:59:51  
 Grid Infrastructure version: 12.1.0.2.00     [PASSED]  
 2015-05-12 14:59:51  
 Checking for active ASM operations....  
 2015-05-12 14:59:51  
 Check for no active ASM operations:       [PASSED]  
 2015-05-12 14:59:51  
 Checking griddisk status across all cells....  
 2015-05-12 14:59:54  
 All griddisks across all cells have asmdeactivationoutcome = Yes  
 All griddisks across all cells are ONLINE  
 Griddisk checks:                 [PASSED]  
 2015-05-12 14:59:54  
 Checking flash cache status.....  
 2015-05-12 14:59:55  
 Flashcache status normal:            [PASSED]  
 2015-05-12 14:59:55  
 Checking that all FlashDisks are present...  
 2015-05-12 14:59:57  
 FlashDisk validation:              [PASSED]  
 2015-05-12 14:59:57  
 Checking current flash cache mode.....  
 2015-05-12 14:59:58  
 Flashcache not already in target mode:      [PASSED]  
 2015-05-12 14:59:58  
 All pre-req checks completed:          [PASSED]  
 2015-05-12 14:59:59  
 dummy01celadm01: flashcache size: 5.82122802734375T  
 dummy01celadm02: flashcache size: 5.82122802734375T  
 dummy01celadm03: flashcache size: 5.82122802734375T  
 dummy01celadm04: flashcache size: 5.82122802734375T  
 dummy01celadm05: flashcache size: 5.82122802734375T  
 dummy01celadm06: flashcache size: 5.82122802734375T  
 dummy01celadm07: flashcache size: 5.82122802734375T  
 There are 7 storage cells to process.  
 2015-05-12 14:59:59  
 Changing flash cache to WriteBack ROLLING....  
 2015-05-12 14:59:59  
 STEP 0: Checking gridisk status on cell: dummy01celadm01  
 2015-05-12 15:00:01  
 STEP 0 completed successfully on cell: dummy01celadm01  
 2015-05-12 15:00:04  
 STEP 1: Dropping flashcache on cell: dummy01celadm01  
 2015-05-12 15:00:36  
 STEP 1: Completed sucessfully on cell: dummy01celadm01  
 2015-05-12 15:00:39  
 Skipping STEP 2: Inactivating grid disks not required.  
 2015-05-12 15:00:39  
 Skipping STEP 3: Shutdown of cellsrv not required.  
 2015-05-12 15:00:39  
 STEP 4: Set the flashCachMode on cell: dummy01celadm01  
 2015-05-12 15:00:39  
 STEP 4: Setting flashCacheMode to WriteBack  
 2015-05-12 15:00:40  
 STEP 4: Completed successfully on cell: dummy01celadm01  
 2015-05-12 15:00:43  
 Skipping STEP 5: Restart of cellsrv not required.  
 2015-05-12 15:00:43  
 Skipping STEP 6: Acativating grid disks not required.  
 2015-05-12 15:00:43  
 STEP 7: Creating flashcache on cell: dummy01celadm01  
 2015-05-12 15:01:32  
 STEP 7: Completed sucessfully on cell: dummy01celadm01  
 2015-05-12 15:01:35  
 STEP 8: Verifying flashCacheMode on cell: dummy01celadm01  
 2015-05-12 15:01:36  
 STEP 8: Completed sucessfully on cell: dummy01celadm01  
 Flash Cache mode is now WriteBack  
 2015-05-12 15:01:39  
 Skipping STEP 9: Waiting for grid disks to sync not required.  
 2015-05-12 15:01:39  
 STEP 0: Checking gridisk status on cell: dummy01celadm02  
 2015-05-12 15:01:41  
 STEP 0 completed successfully on cell: dummy01celadm02  
 2015-05-12 15:01:44  
 STEP 1: Dropping flashcache on cell: dummy01celadm02  
 2015-05-12 15:02:19  
 STEP 1: Completed sucessfully on cell: dummy01celadm02  
 2015-05-12 15:02:22  
 Skipping STEP 2: Inactivating grid disks not required.  
 2015-05-12 15:02:22  
 Skipping STEP 3: Shutdown of cellsrv not required.  
 2015-05-12 15:02:22  
 STEP 4: Set the flashCachMode on cell: dummy01celadm02  
 2015-05-12 15:02:22  
 STEP 4: Setting flashCacheMode to WriteBack  
 2015-05-12 15:02:23  
 STEP 4: Completed successfully on cell: dummy01celadm02  
 2015-05-12 15:02:26  
 Skipping STEP 5: Restart of cellsrv not required.  
 2015-05-12 15:02:26  
 Skipping STEP 6: Acativating grid disks not required.  
 2015-05-12 15:02:26  
 STEP 7: Creating flashcache on cell: dummy01celadm02  
 2015-05-12 15:03:21  
 STEP 7: Completed sucessfully on cell: dummy01celadm02  
 2015-05-12 15:03:24  
 STEP 8: Verifying flashCacheMode on cell: dummy01celadm02  
 2015-05-12 15:03:25  
 STEP 8: Completed sucessfully on cell: dummy01celadm02  
 Flash Cache mode is now WriteBack  
 2015-05-12 15:03:28  
 Skipping STEP 9: Waiting for grid disks to sync not required.  
 2015-05-12 15:03:28  
 STEP 0: Checking gridisk status on cell: dummy01celadm03  
 2015-05-12 15:03:29  
 STEP 0 completed successfully on cell: dummy01celadm03  
 2015-05-12 15:03:32  
 STEP 1: Dropping flashcache on cell: dummy01celadm03  
 2015-05-12 15:04:07  
 STEP 1: Completed sucessfully on cell: dummy01celadm03  
 2015-05-12 15:04:10  
 Skipping STEP 2: Inactivating grid disks not required.  
 2015-05-12 15:04:10  
 Skipping STEP 3: Shutdown of cellsrv not required.  
 2015-05-12 15:04:10  
 STEP 4: Set the flashCachMode on cell: dummy01celadm03  
 2015-05-12 15:04:10  
 STEP 4: Setting flashCacheMode to WriteBack  
 2015-05-12 15:04:10  
 STEP 4: Completed successfully on cell: dummy01celadm03  
 2015-05-12 15:04:13  
 Skipping STEP 5: Restart of cellsrv not required.  
 2015-05-12 15:04:13  
 Skipping STEP 6: Acativating grid disks not required.  
 2015-05-12 15:04:13  
 STEP 7: Creating flashcache on cell: dummy01celadm03  
 2015-05-12 15:05:09  
 STEP 7: Completed sucessfully on cell: dummy01celadm03  
 2015-05-12 15:05:12  
 STEP 8: Verifying flashCacheMode on cell: dummy01celadm03  
 2015-05-12 15:05:12  
 STEP 8: Completed sucessfully on cell: dummy01celadm03  
 Flash Cache mode is now WriteBack  
 2015-05-12 15:05:15  
 Skipping STEP 9: Waiting for grid disks to sync not required.  
 2015-05-12 15:05:15  
 STEP 0: Checking gridisk status on cell: dummy01celadm04  
 2015-05-12 15:05:17  
 STEP 0 completed successfully on cell: dummy01celadm04  
 2015-05-12 15:05:20  
 STEP 1: Dropping flashcache on cell: dummy01celadm04  
 2015-05-12 15:05:53  
 STEP 1: Completed sucessfully on cell: dummy01celadm04  
 2015-05-12 15:05:56  
 Skipping STEP 2: Inactivating grid disks not required.  
 2015-05-12 15:05:56  
 Skipping STEP 3: Shutdown of cellsrv not required.  
 2015-05-12 15:05:56  
 STEP 4: Set the flashCachMode on cell: dummy01celadm04  
 2015-05-12 15:05:56  
 STEP 4: Setting flashCacheMode to WriteBack  
 2015-05-12 15:06:01  
 STEP 4: Completed successfully on cell: dummy01celadm04  
 2015-05-12 15:06:04  
 Skipping STEP 5: Restart of cellsrv not required.  
 2015-05-12 15:06:04  
 Skipping STEP 6: Acativating grid disks not required.  
 2015-05-12 15:06:04  
 STEP 7: Creating flashcache on cell: dummy01celadm04  
 2015-05-12 15:06:47  
 STEP 7: Completed sucessfully on cell: dummy01celadm04  
 2015-05-12 15:06:50  
 STEP 8: Verifying flashCacheMode on cell: dummy01celadm04  
 2015-05-12 15:06:51  
 STEP 8: Completed sucessfully on cell: dummy01celadm04  
 Flash Cache mode is now WriteBack  
 2015-05-12 15:06:54  
 Skipping STEP 9: Waiting for grid disks to sync not required.  
 2015-05-12 15:06:54  
 STEP 0: Checking gridisk status on cell: dummy01celadm05  
 2015-05-12 15:06:55  
 STEP 0 completed successfully on cell: dummy01celadm05  
 2015-05-12 15:06:58  
 STEP 1: Dropping flashcache on cell: dummy01celadm05  
 2015-05-12 15:07:40  
 STEP 1: Completed sucessfully on cell: dummy01celadm05  
 2015-05-12 15:07:43  
 Skipping STEP 2: Inactivating grid disks not required.  
 2015-05-12 15:07:43  
 Skipping STEP 3: Shutdown of cellsrv not required.  
 2015-05-12 15:07:43  
 STEP 4: Set the flashCachMode on cell: dummy01celadm05  
 2015-05-12 15:07:43  
 STEP 4: Setting flashCacheMode to WriteBack  
 2015-05-12 15:07:43  
 STEP 4: Completed successfully on cell: dummy01celadm05  
 2015-05-12 15:07:46  
 Skipping STEP 5: Restart of cellsrv not required.  
 2015-05-12 15:07:46  
 Skipping STEP 6: Acativating grid disks not required.  
 2015-05-12 15:07:46  
 STEP 7: Creating flashcache on cell: dummy01celadm05  
 2015-05-12 15:08:37  
 STEP 7: Completed sucessfully on cell: dummy01celadm05  
 2015-05-12 15:08:40  
 STEP 8: Verifying flashCacheMode on cell: dummy01celadm05  
 2015-05-12 15:08:41  
 STEP 8: Completed sucessfully on cell: dummy01celadm05  
 Flash Cache mode is now WriteBack  
 2015-05-12 15:08:44  
 Skipping STEP 9: Waiting for grid disks to sync not required.  
 2015-05-12 15:08:44  
 STEP 0: Checking gridisk status on cell: dummy01celadm06  
 2015-05-12 15:08:46  
 STEP 0 completed successfully on cell: dummy01celadm06  
 2015-05-12 15:08:49  
 STEP 1: Dropping flashcache on cell: dummy01celadm06  
 2015-05-12 15:09:23  
 STEP 1: Completed sucessfully on cell: dummy01celadm06  
 2015-05-12 15:09:26  
 Skipping STEP 2: Inactivating grid disks not required.  
 2015-05-12 15:09:26  
 Skipping STEP 3: Shutdown of cellsrv not required.  
 2015-05-12 15:09:26  
 STEP 4: Set the flashCachMode on cell: dummy01celadm06  
 2015-05-12 15:09:26  
 STEP 4: Setting flashCacheMode to WriteBack  
 2015-05-12 15:09:30  
 STEP 4: Completed successfully on cell: dummy01celadm06  
 2015-05-12 15:09:33  
 Skipping STEP 5: Restart of cellsrv not required.  
 2015-05-12 15:09:33  
 Skipping STEP 6: Acativating grid disks not required.  
 2015-05-12 15:09:33  
 STEP 7: Creating flashcache on cell: dummy01celadm06  
 2015-05-12 15:10:19  
 STEP 7: Completed sucessfully on cell: dummy01celadm06  
 2015-05-12 15:10:22  
 STEP 8: Verifying flashCacheMode on cell: dummy01celadm06  
 2015-05-12 15:10:23  
 STEP 8: Completed sucessfully on cell: dummy01celadm06  
 Flash Cache mode is now WriteBack  
 2015-05-12 15:10:26  
 Skipping STEP 9: Waiting for grid disks to sync not required.  
 2015-05-12 15:10:26  
 STEP 0: Checking gridisk status on cell: dummy01celadm07  
 2015-05-12 15:10:28  
 STEP 0 completed successfully on cell: dummy01celadm07  
 2015-05-12 15:10:31  
 STEP 1: Dropping flashcache on cell: dummy01celadm07  
 2015-05-12 15:11:04  
 STEP 1: Completed sucessfully on cell: dummy01celadm07  
 2015-05-12 15:11:07  
 Skipping STEP 2: Inactivating grid disks not required.  
 2015-05-12 15:11:07  
 Skipping STEP 3: Shutdown of cellsrv not required.  
 2015-05-12 15:11:07  
 STEP 4: Set the flashCachMode on cell: dummy01celadm07  
 2015-05-12 15:11:07  
 STEP 4: Setting flashCacheMode to WriteBack  
 2015-05-12 15:11:13  
 STEP 4: Completed successfully on cell: dummy01celadm07  
 2015-05-12 15:11:16  
 Skipping STEP 5: Restart of cellsrv not required.  
 2015-05-12 15:11:16  
 Skipping STEP 6: Acativating grid disks not required.  
 2015-05-12 15:11:16  
 STEP 7: Creating flashcache on cell: dummy01celadm07  
 2015-05-12 15:12:02  
 STEP 7: Completed sucessfully on cell: dummy01celadm07  
 2015-05-12 15:12:05  
 STEP 8: Verifying flashCacheMode on cell: dummy01celadm07  
 2015-05-12 15:12:06  
 STEP 8: Completed sucessfully on cell: dummy01celadm07  
 Flash Cache mode is now WriteBack  
 2015-05-12 15:12:09  
 Skipping STEP 9: Waiting for grid disks to sync not required.  
 2015-05-12 15:12:09  
 Validating inventory for griddisks  
 2015-05-12 15:12:11  
 Validation of griddisk:             [PASSED]  
 2015-05-12 15:12:11  
 Validating inventory for flashdisks  
 2015-05-12 15:12:13  
 Validation of flashdisk:             [PASSED]  
 2015-05-12 15:12:13  
 Validating inventory for flashsize  
 2015-05-12 15:12:14  
 Validation of flashsize:             [PASSED]  
 2015-05-12 15:12:14  
 Setting flash cache to WriteBack completed successfully.  

dcli -g ~/cell_group -l root cellcli -e "list cell attributes flashcachemode"

Tuesday, May 5, 2015

Exadata Local Read-only file system | End_request: I/O error

---On database Nodes we wont be able to write on / or any local file system.

 [root@dummyhostname01 ~]# df  
 Filesystem       1K-blocks    Used  Available Use% Mounted on  
 /dev/mapper/VGExaDb-LVDbSys1  
             30963708  22372264   7018580 77% /  
 tmpfs         264152064     4  264152060  1% /dev/shm  
 /dev/sda1         516040   40016   449812  9% /boot  
 /dev/mapper/VGExaDb-LVDbOra1  
             103212320  25806260  72163180 27% /u01  
 [root@dummyhostname01 oracle.cellos]# cd conf  
 [root@dummyhostname01 conf]# ls  
 ls: reading directory .: Input/output error  
 [root@dummyhostname01 log]# cd /u01  
 [root@dummyhostname01 u01]# ls  
 app lost+found  
 [root@dummyhostname01 u01]# ll  
 total 20  
 drwxr-xr-x 5 root oinstall 4096 Mar 23 19:11 app  
 drwx------ 2 root root   16384 Mar 10 19:26 lost+found  
 [root@dummyhostname01 u01]# touch test  
 touch: cannot touch `test': Read-only file system  


---On looking at tail -f /var/log/messeges

 May 5 02:18:24 dummyhostname01 adclient[21135]: INFO <fd:25 sudo(100222)> client.sudo Set credentials for user 'root': mapping misconfiguration. Passing user to next service module.  
 May 5 02:18:26 dummyhostname01 kernel: megaraid_sas: Iop2SysDoorbellIntfor scsi0  
 May 5 02:18:27 dummyhostname01 kernel: megasas: Found FW in FAULT state, will reset adapter scsi0.  
 May 5 02:18:27 dummyhostname01 kernel: megaraid_sas: resetting fusion adapter scsi0.  
 May 5 02:18:27 dummyhostname01 kernel: megaraid_sas: Reset not supported, killing adapter scsi0.  
 May 5 02:18:27 dummyhostname01 kernel: sd 0:2:0:0: [sda] Unhandled error code  
 May 5 02:18:27 dummyhostname01 kernel: sd 0:2:0:0: [sda] Result: hostbyte=DID_NO_CONNECT driverbyte=DRIVER_OK  
 May 5 02:18:27 dummyhostname01 kernel: sd 0:2:0:0: [sda] CDB: Write(10): 2a 00 16 b5 5d 40 00 00 08 00  
 May 5 02:18:27 dummyhostname01 kernel: blk_update_request: 4 callbacks suppressed  
 May 5 02:18:27 dummyhostname01 kernel: end_request: I/O error, dev sda, sector 380984640  
 May 5 02:18:27 dummyhostname01 kernel: Buffer I/O error on device dm-3, logical block 17606304  

---firmware / hardware diagnostic on Database Node.
---As you can not write on Local Filesystem you could write on NFS mounts if there is present on DBNODE.

 dmesg > /NFS_MOUNT/sundiag/dmesg.txt   
 ipmitool sunoem cli force 'show /SP/console/history' > /NFS_MOUNT/sundiag/console.out   
 /opt/MegaRAID/MegaCli/MegaCli64 -FwTermLog Dsply -a0 >/NFS_MOUNT/sundiag/fwterm.txt   
 [root@dummyhostname01 sundiag]# cat /NFS_MOUNT/sundiag/fwterm.txt  
 User specified controller is not present.  
 Failed to get CpController object.  
 Exit Code: 0x01  
 /opt/MegaRAID/MegaCli/MegaCli64 -AdpEventLog -GetEvents -f /NFS_MOUNT/sundiag/events.txt -a0   
 lspci -vvv > /NFS_MOUNT/sundiag/lspci.out  

Solution.

--Finally reboot / restart force fully from ILOM.

 -> stop -f /SYS  
 Are you sure you want to immediately stop /SYS (y/n)? y  
 Stopping /SYS immediately  

Wednesday, February 25, 2015

RS-700 Celloflsrv hang detected on Exadata 12.1.1.1

RS-700 [Celloflsrv hang detected. It will be terminated] [SYS_121111_140712] [] [] [] [] [] [] [] [] [] []
Server Model Oracle Corporation SUN SERVER X4-2L High Capacity
Release Version 12.1.1.1.1
Release Label OSS_12.1.1.1.1_LINUX.X64_140712

This is Bug 19132065 - Oracle Linux semtimedop() wakeups by timeout are lagging causing offload operations to fail (which may degrade performance) and errors similar to one or more of the following:
? ORA-700 [Offload issue job timed out]
? ORA-700 [Offload group not open]
? RS-700 [Celloflsrv hang detected. It will be terminated]

This bug affects related to 12.1.1.1. storage Version.
It is due to DB Node RCU delayed and cause Offload job to fail on Cellservices .
it affects database performance not availability.
Error ocure mostly when cellserv tried to do Read optimization.
reducing Delay in RCU is work around accross whole stack.

Step 1: Set rcu_delay for runtime

# echo 1 > /proc/sys/kernel/rcu_delay
Verify the setting
# cat /proc/sys/kernel/rcu_delay
1

Step 2: Set rcu_delay in /etc/sysctl.conf for proper setting upon reboot

Add "kernel.rcu_delay=1" to /etc/sysctl.conf

Step 3: Restart cellsrv on storage servers

CellCLI> alter cell restart services cellsrv;


This workaround is automatically applied in the following cases:
When a new system is deployed with Exadata 11.2.3.3.1 or 12.1.1.1.1 using OEDA Sep 2014 or later.
When storage servers are upgraded to 11.2.3.3.1 or 12.1.1.1.1 and the patchmgr plugins patch is properly staged before running patchmgr, as documented.
When database servers are upgraded to 11.2.3.3.1 or 12.1.1.1.1 using dbnodeupdate.sh v3.58 or later.

References https://www.kernel.org/doc/Documentation/RCU/whatisRCU.txt

Tuesday, November 18, 2014

Exadata Schedule Exachk using OEM and command Line | Exadata healthcheck OEM

There is New Plugin to Install and schedule Exachk on Exadata/ Exalogic system Through 12c Grid control.

Plugin has to first deploy on OMS server and then Exadata / Exalogic Agents.

12c Home page setup (right upper corner) -> extensibility -> Plugin -> Engineered system -> Oracle Exadata healthcheck.

Once it deploy on 12c Grid OMS and Exadata agent.
You can create Monitoring Targets for Exadata Healthcheck Metrics.

https://docs.oracle.com/cd/E24628_01/install.121/e27420/toc.htm#PICHK105 

To schedule Using command Line.

Download Latest Exachk from
Oracle Exadata Database Machine exachk or HealthCheck (Doc ID 1070954.1).
Exachk Health-Check Tool for Exalogic (Doc ID 1449226.1)


Note : Exachk must run from Enterprise Vserver in Virtual Exalogic Environment.

---check Version

./exachk -v
EXACHK  VERSION: 12.1.0.2.1

---Set User equivelency between DB node to storage cell and IB switches 

./exachk -initpresetup   

---setup auto restart and Exachk Deamon. 

./exachk -initsetup
Setting up exachk auto restart functionality using inittab
Starting exachk daemon. . .  .

---Check auto restart and deamon status . 

./exachk -initcheck
Auto restart functionality is configured.
exachk daemon is running. PID : 00000

---schedule exachk. 

AUTORUN_SCHEDULE * * * *       :- Automatic run at specific time in daemon mode.
                 - - - -
                 ? ? ? ?
                 ? ? ? +----- day of week (0 - 6) (0 to 6 are Sunday to Saturday)
                 ? ? +---------- month (1 - 12)
                 ? +--------------- day of month (1 - 31)
                 +-------------------- hour (0 - 23)

./exachk -set "AUTORUN_SCHEDULE=23 * * 0;AUTORUN_FLAGS= -a -o v;NOTIFICATION_EMAIL=firstname.lastname@company.com;PASSWORD_CHECK_INTERVAL=1;"

---List all parameter for exachk. 

./exachk -get all

Sunday, November 2, 2014

Exadata Patching | Upgrade Exadata

This is High Level step by step Instruction for APR-2014 (11.2.0.4) Exadata QFSDP.

- if you are at Jan-2105 . do not place PSU into NFS mount.
- do not use -s to shutdown cluster if you are using dbnodeupdate.sh 4.13
- You must force_reset everytime you run patch precheck.
- dbnodeupdate.sh will run only on Node you are patching. reverse of Upgrading Exalogic compute Nodes.
- patchmgr always run from database nodes.


It is divided in Three major Part.

-Upgrade Exadata Database server RPMs
-Upgrade Storage server Image and Infiniband Switch.
-Upgrade Grid Home and Oracle Home .

1 ExadataDatabaseServer
This Part required Reboot and shutdown of CRS on Local Node. We will apply This patch In rolling.
Database Node gets updates from script provide along with Exadata QFSDP zip file.
Check usage of script.
dbnodeupdate.sh: Exadata Database Server Patching using the DB Node Update Utility (Doc ID 1553103.1)

1.1 copy zip file to Local directory and Inflate dbnodeupdate script comes with QFSDP

--You can skip this part if patch is located on Local storage.

 dcli -g db_group -l root "mkdir /u01/patches/YUM/”  
 dcli -g db_group -l root "cp <patch unzip location>/18370227/Infrastructure/11.2.3.3.0/ExadataDatabaseServer/p17809253_112330_Linux-x86-64.zip /u01/patches/YUM"  

This will inflate dbnodeupdate.sh script.

 cd <patch_location>/18370227/Infrastructure/ExadataDBNodeUpdate/3.2  
 unzip p16486998_121110_Linux-x86-64.zip  

Usage of script

 Usage: dbnodeupdate.sh [ -u | -r | -c ] -l <baseurl|zip file> [-p] <phase> [-n] [-s] [-q] [-v] [-t] [-a] <alert.sh> [-b] [-m] | [-V] | [-h]  
 -u            Upgrade  
 -r            Rollback  
 -c            Complete post actions (relink all homes, enable GI to start)  
 -l <baseurl|zip file>  Baseurl (http or zipped iso file for the repository)  
 -s            Shutdown stack before upgrading/rolling back  
 -p            Bootstrap phase (1 or 2) only to be used when instructed by dbnodeupdate.sh  
 -q            Quiet mode (no prompting) only be used in combination with -t  
 -n            No backup will be created  
 -t            'to release' - used when in quiet mode or used when updating to one-offs/releases via 'latest' channel (requires 11.2.3.2.1)  
 -v            Verify prereqs only. Only to be used with -u and -l option  
 -b            Peform backup only  
 -a <alert.sh>      Full path to shell script used for alert trapping  
 -m            Install / update-to exadata-sun/hp-computenode-minimum only (11.2.3.3.0 and later)  
 -V            Print version  
 -h            Print usage  

1.2 Run pre-upgrade steps .

 ./dbnodeupdate.sh -u -l /u01/patches/YUM/p18876946_112331_Linux-x86-64.zip -v  
 ##########################################################################################################################  
 #                                                                           #  
 # Guidelines for using dbnodeupdate.sh (rel. 3.53):                                            #  
 #                                                            #  
 # - Prerequisites for usage:                                               #  
 #     1. Refer to dbnodeupdate.sh options. See MOS 1553103.1                             #  
 #     2. Use the latest release of dbnodeupdate.sh. See patch 16486998                        #  
 #     3. Run the prereq check with the '-v' option.                                 #  
 #                                                            #  
 #  I.e.: ./dbnodeupdate.sh -u -l /u01/my-iso-repo.zip -v                                #  
 #     ./dbnodeupdate.sh -u -l http://my-yum-repo -v                                 #  
 #                                                            #  
 # - Prerequisite dependency check failures can happen due to customization:                       #  
 #   - The prereq check detects dependency issues that need to be addressed prior to running a successful update.    #  
 #   - Customized rpm packages may fail the built-in dependency check and system updates cannot proceed until resolved. #  
 #                                                            #  
 #  When upgrading from releases later than 11.2.2.4.2 to releases before 11.2.3.3.0:                  #  
 #   - Conflicting packages should be removed before proceeding the update.                      #  
 #                                                            #  
 #  When upgrading to releases 11.2.3.3.0 or later:                                   #  
 #   - When the 'exact' package dependency check fails 'minimum' package dependency check will be tried.        #  
 #   - When the 'minimum' package dependency check also fails,                             #  
 #    the conflicting packages should be removed before proceeding.                          #  
 #                                                            #  
 # - As part of the prereq checks and as part of the update, a number of rpms will be removed.              #  
 #  This removal is required to preserve Exadata functioning. This should not be confused with obsolete packages.    #  
 #   - See /var/log/cellos/packages_to_be_removed.txt for details on what packages will be removed.          #  
 #                                                            #  
 # - In case of any problem when filing an SR, upload the following:                           #  
 #   - /var/log/cellos/dbnodeupdate.log                                        #  
 #   - /var/log/cellos/dbnodeupdate.<runid>.diag                                    #  
 #   - where <runid> is the unique number of the failing run.                             #  
 #                                                            #  
 ##########################################################################################################################  
 Continue ? [y/n]  
 y  
  (*) 2015-02-01 15:46:28: Unzipping helpers (QFSDP_JULY2014_EXADATA/19069261/Infrastructure/ExadataDBNodeUpdate/3.53/dbupdate-helpers.zip) to /opt/oracle.SupportTools/dbnodeupdate_helpers  
  (*) 2015-02-01 15:46:29: Initializing logfile /var/log/cellos/dbnodeupdate.log  
  Warning: Active NFS and/or SMBFS mounts found on this DB node.  
       Before taking a backup or performing the actual update these need to be unmounted.  
       For the actual update (not now) dbnodeupdate.sh will try unmounting them silently.  
       During collection of system configuration (prereq) stale network mounts may cause long waits and dbnodeupdate.sh to stall  
       It is therefore recommended (not required) to unmount any active network mount now before continuing.  
 Continue ? [y/n]  
 y  
  (*) 2015-02-01 15:47:52: Collecting system configuration details. This may take a while...  
  (*) 2015-02-01 15:48:40: Validating system details for known issues and best practices. This may take a while...  
  (*) 2015-02-01 15:48:40: Checking free space in /u01/patches/YUM/iso.stage.010215154537  
  (*) 2015-02-01 15:48:40: Unzipping /u01/patches/YUM/p18876946_112331_Linux-x86-64.zip to /u01/patches/YUM/iso.stage.010215154537, this may take a while  
  (*) 2015-02-01 15:48:51: Original /etc/yum.conf moved to /etc/yum.conf.010215154537, generating new yum.conf  
  (*) 2015-02-01 15:48:51: Generating Exadata repository file /etc/yum.repos.d/Exadata-computenode.repo  
  ERROR: Duplicate entries detected in /etc/fstab. Correct settings and rerun dbnodeupdate.sh.  
  (*) 2015-02-01 15:50:03: Cleaning up iso and temp mount points  
 p –v    



1.2 Upgrade Compute Node From Local Yum zip file.

This command upgrade / and reboot Compute Node at end Of Patch. If we include –s it will stop CRS in this Node.
you can monitor progress by tail -f /var/log/cellos/dbnodeupdate.log

 ./dbnodeupdate.sh -u -l /u01/patches/YUM/p18876946_112331_Linux-x86-64.zip -n -s  
  (*) 2015-02-03 00:12:29: Cleaning up the yum cache.  
  (*) 2015-02-03 00:12:31: Performing yum package dependency check for 'exact' dependencies. This may take a while...  
  (*) 2015-02-03 00:12:33: 'Exact' package dependency check failed.  
  (*) 2015-02-03 00:12:54: Performing yum package dependency check for 'minimum' dependencies. This may take a while...  
  (*) 2015-02-03 00:12:56: 'Minimum' package dependency check succeeded.  
 Active Image version  : 11.2.3.3.0.131014.1  
 Active Kernel version : 2.6.39-400.126.1.el5uek  
 Active LVM Name    : /dev/mapper/VGExaDb-LVDbSys1  
 Inactive Image version : n/a  
 Inactive LVM Name   : /dev/mapper/VGExaDb-LVDbSys2  
 Current user id    : root  
 Action         : upgrade  
 Upgrading to      : 11.2.3.3.1.140529.1 (to exadata-sun-computenode-minimum)  
 Baseurl        : file:///var/www/html/yum/unknown/EXADATA/dbserver/030215001041/x86_64/ (iso)  
 Iso file        : /u01/patches/YUM/iso.stage.030215001041/112331_base_repo_140529.1.iso  
 Create a backup    : No  
 Shutdown stack     : Yes (Currently stack is up)  
 RPM exclusion list   : Not in use (add rpms to /etc/exadata/yum/exclusion.lst and restart dbnodeupdate.sh)  
 RPM obsolete list   : /etc/exadata/yum/obsolete.lst (lists rpms to be removed by the update)  
             : RPM obsolete list is extracted from exadata-sun-computenode-11.2.3.3.1.140529.1-1.x86_64.rpm  
 Exact dependencies   : Will fail on a next update. Update to 'exact' will be not possible. Falling back to 'minimum'  
             : See /var/log/cellos/exact_conflict_report.030215001041.txt for more details  
             : Update target switched to 'minimum'  
 Minimum dependencies  : No conflicts  
 Logfile        : /var/log/cellos/dbnodeupdate.log (runid: 030215001041)  
 Diagfile        : /var/log/cellos/dbnodeupdate.030215001041.diag  
 Server model      : SUN SERVER X4-2  
 Remote mounts exist  : Yes (dbnodeupdate.sh will try unmounting)  
 dbnodeupdate.sh rel.  : 3.53 (always check MOS 1553103.1 for the latest release of dbnodeupdate)  
 Note          : After upgrading and rebooting run './dbnodeupdate.sh -c' to finish post steps.  
 The following known issues will be checked for but require manual follow-up:  
  (*) - Issue - Yum rolling update requires fix for 11768055 when Grid Infrastructure is below 11.2.0.2 BP12  
 Continue ? [y/n]  
 Continue ? [y/n]  
 y  
  (*) 2015-02-03 00:13:06: Verifying GI and DB's are shutdown  
  (*) 2015-02-03 00:13:06: Shutting down GI and db  
  (*) 2015-02-03 00:13:36: Collecting console history for diag purposes  
  (*) 2015-02-03 00:14:00: Successfully unmounted network mount /nfs_mount/backup02  
  (*) 2015-02-03 00:14:00: Successfully unmounted network mount /nfs_mount/backup01  
  (*) 2015-02-03 00:14:05: Successfully unmounted network mount /nfs_mount/backup01  
  (*) 2015-02-03 00:14:06: Successfully unmounted network mount /nfs_mount/backup02  
  (*) 2015-02-03 00:14:06: Successfully unmounted network mount /nfs_mount/p01_bak01  
  (*) 2015-02-03 00:14:06: Successfully unmounted network mount /nfs_mount/p01_bak02  
  (*) 2015-02-03 00:14:06: Successfully unmounted network mount /nfs_mount/p01_bak01  
  (*) 2015-02-03 00:14:06: Successfully unmounted network mount /nfs_mount/p01_bak02  
  (*) 2015-02-03 00:14:06: Unmount of /boot successful  
  (*) 2015-02-03 00:14:06: Check for /dev/sda1 successful  
  (*) 2015-02-03 00:14:06: Mount of /boot successful  
  (*) 2015-02-03 00:14:06: Disabling stack from starting  
  (*) 2015-02-03 00:14:13: ExaWatcher stopped successful  
  (*) 2015-02-03 00:14:13: Validating the specified source location.  
  (*) 2015-02-03 00:14:14: Cleaning up the yum cache.  
  (*) 2015-02-03 00:14:17: Performing yum update. Node is expected to reboot when finished.  
  (*) 2015-02-03 00:16:45: Waiting for post rpm script to finish. Sleeping another 60 seconds (60 / 900)  
 Remote broadcast message (Tue Feb 3 00:16:53 2015):  
 Exadata post install steps started.  
 It may take up to 15 minutes.  
  (*) 2015-02-03 00:17:45: Waiting for post rpm script to finish. Sleeping another 60 seconds (120 / 900)  
  (*) 2015-02-03 00:18:45: Waiting for post rpm script to finish. Sleeping another 60 seconds (180 / 900)  
  (*) 2015-02-03 00:19:45: Waiting for post rpm script to finish. Sleeping another 60 seconds (240 / 900)  
  (*) 2015-02-03 00:20:45: Waiting for post rpm script to finish. Sleeping another 60 seconds (300 / 900)  
  (*) 2015-02-03 00:21:45: Waiting for post rpm script to finish. Sleeping another 60 seconds (360 / 900)  
  (*) 2015-02-03 00:22:45: Waiting for post rpm script to finish. Sleeping another 60 seconds (420 / 900)  
 Remote broadcast message (Tue Feb 3 00:23:13 2015):  
 Exadata post install steps completed.  
  (*) 2015-02-03 00:23:45: Waiting for post rpm script to finish. Sleeping another 60 seconds (480 / 900)  
  (*) 2015-02-03 00:24:46: All post steps are finished.  
  (*) 2015-02-03 00:24:46: System will reboot automatically for changes to take effect  
  (*) 2015-02-03 00:24:46: After reboot run "./dbnodeupdate.sh -c" to complete the upgrade  
  (*) 2015-02-03 00:25:05: Cleaning up iso and temp mount points  
  (*) 2015-02-03 00:25:06: Rebooting now...  
 Broadcast message from root (pts/6) (Tue Feb 3 00:25:06 2015):  
 The system is going down for reboot NOW!  
 ----------------------------  
 1st time reboot.   
 ----------------------------  
 ./dbnodeupdate.sh -c  
 Continue ? [y/n]  
 y  
  (*) 2015-02-03 01:49:49: Unzipping helpers (/QFSDP_JULY2014_EXADATA/19069261/Infrastructure/ExadataDBNodeUpdate/3.53/dbupdate-helpers.zip) to /opt/oracle.SupportTools/dbnodeupdate_helpers  
  (*) 2015-02-03 01:49:49: Initializing logfile /var/log/cellos/dbnodeupdate.log  
  (*) 2015-02-03 01:49:50: Collecting system configuration details. This may take a while...  
 Active Image version  : 11.2.3.3.1.140529.1  
 Active Kernel version : 2.6.39-400.128.17.el5uek  
 Active LVM Name    : /dev/mapper/VGExaDb-LVDbSys1  
 Inactive Image version : n/a  
 Inactive LVM Name   : /dev/mapper/VGExaDb-LVDbSys2  
 Current user id    : root  
 Action         : finish-post (validate image status, fix known issues, cleanup, relink and enable crs to auto-start)  
 Shutdown stack     : No (Currently stack is down)  
 Logfile        : /var/log/cellos/dbnodeupdate.log (runid: 030215014947)  
 Diagfile        : /var/log/cellos/dbnodeupdate.030215014947.diag  
 Server model      : SUN SERVER X4-2  
 dbnodeupdate.sh rel.  : 3.53 (always check MOS 1553103.1 for the latest release of dbnodeupdate)  
 The following known issues will be checked for but require manual follow-up:  
  (*) - Issue - Yum rolling update requires fix for 11768055 when Grid Infrastructure is below 11.2.0.2 BP12  
 Continue ? [y/n]  
 y  
  (*) 2015-02-03 01:54:28: Verifying GI and DB's are shutdown  
  (*) 2015-02-03 01:54:31: Verifying firmware updates/validations. Maximum wait time: 60 minutes.  
  (*) 2015-02-03 01:54:31: If the node reboots during this firmware update/validation, re-run './dbnodeupdate.sh -c' after the node restarts.........  
 Broadcast message from root (console) (Tue Feb 3 02:03:08 2015):  
 The system is going down for system halt NOW!  
 ----------------------------  
 2nd time reboot.   
 ----------------------------  
 [root@pwerxd01dbadm04 3.53]# ./dbnodeupdate.sh -c  
 Continue ? [y/n]  
 y  
  (*) 2015-02-03 02:13:44: Unzipping helpers (/19069261/Infrastructure/ExadataDBNodeUpdate/3.53/dbupdate-helpers.zip) to /opt/oracle.SupportTools/dbnodeupdate_helpers  
  (*) 2015-02-03 02:13:45: Initializing logfile /var/log/cellos/dbnodeupdate.log  
  (*) 2015-02-03 02:13:45: Collecting system configuration details. This may take a while...  
 Active Image version  : 11.2.3.3.1.140529.1  
 Active Kernel version : 2.6.39-400.128.17.el5uek  
 Active LVM Name    : /dev/mapper/VGExaDb-LVDbSys1  
 Inactive Image version : n/a  
 Inactive LVM Name   : /dev/mapper/VGExaDb-LVDbSys2  
 Current user id    : root  
 Action         : finish-post (validate image status, fix known issues, cleanup, relink and enable crs to auto-start)  
 Shutdown stack     : No (Currently stack is down)  
 Logfile        : /var/log/cellos/dbnodeupdate.log (runid: 030215021342)  
 Diagfile        : /var/log/cellos/dbnodeupdate.030215021342.diag  
 Server model      : SUN SERVER X4-2  
 dbnodeupdate.sh rel.  : 3.53 (always check MOS 1553103.1 for the latest release of dbnodeupdate)  
 The following known issues will be checked for but require manual follow-up:  
  (*) - Issue - Yum rolling update requires fix for 11768055 when Grid Infrastructure is below 11.2.0.2 BP12  
 Continue ? [y/n]  
 y  
  (*) 2015-02-03 01:16:00: Verifying GI and DB's are shutdown  
  (*) 2015-02-03 01:16:02: Verifying firmware updates/validations. Maximum wait time: 60 minutes.  
  (*) 2015-02-03 01:16:02: If the node reboots during this firmware update/validation, re-run './dbnodeupdate.sh -c' after the node restarts..  
  (*) 2015-02-03 01:16:02: Collecting console history for diag purposes  
  (*) 2015-02-03 01:16:23: No rpms to remove  
  (*) 2015-02-03 01:16:45: EM Agent (in /u01/app/EMbase/core/12.1.0.3.0) stopped successfully  
  (*) 2015-02-03 01:16:45: Relinking all homes  
  (*) 2015-02-03 01:16:45: Unlocking /u01/app/11.2.0.4/grid  
  (*) 2015-02-03 01:16:51: Relinking /oracle/product/11.2.0.3 as orapnacod04 (WARNING: this home is not linked with rds - relink will also be done without rds option)  
  (*) 2015-02-03 01:17:02: Relinking /oracle/product/11.2.0.3 as orapnacop01 (with rds option)  
  (*) 2015-02-03 01:17:14: Relinking /oracle/product/11.2.0.3 as orapnacoq01 (WARNING: this home is not linked with rds - relink will also be done without rds option)  
  (*) 2015-02-03 01:17:25: Relinking /u01/app/11.2.0.4/grid as grid (with rds option)  
  (*) 2015-02-03 01:17:38: Executing /u01/app/11.2.0.4/grid/crs/install/rootcrs.pl -patch  
  (*) 2015-02-03 01:19:44: Sleeping another 60 seconds while stack is starting (1/5)  
  (*) 2015-02-03 01:19:44: Stack started  
  (*) 2015-02-03 01:19:44: Enabling stack to start at reboot. Disable this when the stack should not be starting on a next boot  
  (*) 2015-02-03 01:20:28: EM Agent (in /u01/app/EMbase/core/12.1.0.3.0) started successfully  
  (*) 2015-02-03 01:20:28: All post steps are finished.  

2 ExadataStorageServer and InfiniBandSwitch

2.1 Download Patch manager Plugin from 1487339.1 Metalink Note.

2.2 Preparing Exadata Cells for Patch Application

Execute below command as Root from Compute Node.
Generate rsa/dsa Key on Compute Node.

 ssh-keygen -t rsa  
 ssh-keygen -t dsa  
 Push key to all cell   
 dcli -l root -g cell_group –k  
 dcli -g cell_group -l root 'hostname -i'  

2.3 Set DISK REPAIR TIME on ASM disks.

 select dg.name,a.value from v$asm_diskgroup dg, v$asm_attribute a where dg.group_number=a.group_number and a.name='disk_repair_time';  
 alter diskgroup diskgroup_name set attribute 'disk_repair_time'='3.6h';  

2.4 Turn off DB services on compute Node for NON-ROLLING patch.

NOTE: THIS PART APPLY ONLY IF YOU DECIDE TO APPLY PATCH IN NON-ROLLING FASHION , DO NOT SHUT-DOWN IF PATCH IS ROLLING.

 Run below command as ROOT. From compute Node.   
 dcli -g dbs_group -l root "/u01/app/11.2.0/grid/bin/crsctl stop crs -f"  
 dcli -g dbs_group -l root "ps -ef | grep grid"   
 dcli -g cell_group -l root "cellcli -e alter cell shutdown services all"  

2.5 Pre-check on Patch on Cell storage using patch manager.

 cd <patchlocation>/18370227/Infrastructure/11.2.3.3.0/ExadataStorageServer_InfiniBandSwitch/patch_11.2.3.3.0.131014.1  
 ./patchmgr -cells ~/cell_group -reset_force  
 ./patchmgr -cells cell_group -patch_check_prereq [-rolling] [-ignore_alerts] [-smtp_from "addr" -smtp_to "addr1 addr2 addr3 ..."]  


2.6 Patch cell by Patch Utility.

If the prerequisite checks pass, then start patch application. Use -rolling option if you plan to use rolling updates. Use the -ignore_alerts option to ignore any open hardware alerts on the cells, and continue. Use the -smtp_from, -smtp_to options to set an e-mail address to receive patchmgr alert messages, and continue.

 ./patchmgr -cells ~/cell_group -reset_force  
 ./patchmgr -cells cell_group -patch [-rolling] [-ignore_alerts] [-smtp_from "from_email_address"] [-smtp_to "to_email_address1  to_email_address2 ..."]  

2.7 Check any Grid disk are inactive or offline.

 dcli -g ~/cell_group -l root                     \  
 "cat /root/attempted_deactivated_by_patch_griddisks.txt | grep -v   \  
 ACTIVATE | while read line; do str=\`cellcli -e list griddisk where  \  
 name = \$line attributes name, status, asmmodestatus\`; echo \$str | \  
 grep -v \"active ONLINE\"; done"  

2.8 Run ibswitch pre-requisite

 ./patchmgr –ibswitches -upgrade -ibswitch_precheck   

2.9 Apply patch on IBSWITCH.

 cd <patch location>/18370227/Infrastructure/11.2.3.3.0/ExadataStorageServer_InfiniBandSwitch/patch_11.2.3.3.0.131014.1  
 ./patchmgr –ibswitches -upgrade   

3 Database and Grid Home Upgrade.

3.1 Distribute GI and OH patch to NFS or /tmp
3.1 Install Latest OPatch and Oplan.
3.2 Generate steps to Patch GI using Oplan

 <$GRID_HOME>/OPatch/oplan generateApplySteps <patch location>/18370227/database/11.2.0.4.6_QDPE_Apr2014/18371656  
 <$GRID_HOME>/OPatch/oplan generateRollbackSteps <patch location>/18370227/database/11.2.0.4.6_QDPE_Apr2014/18371656  

3.3 Create OCM file

 dcli -g ~/dbs_group -l oracle $ORACLE_HOME/OPatch/ocm/bin/emocmrsp –output /home/oracle  

3.4 Follow Oplan generated File Instruction.

Wednesday, May 21, 2014

Exadata Bug-list fixed from Image 11.2.3.3 and Later.

As the Exadata X4-2 with 11.2.3.3 and later has below bugs fixed.

Fixes:
-----

9964936     ENHANCE EXADATA ASR TO FILE AUTO SRS FOR (PREDICTIVE) BATTERY FAILURES
11065811     FOR CELLDISK IMPORT REQUIRED PUBLISH EVENT TO ASM + DEL ENDIANNES FRM OWNER FILE
11683510     BATTERY TEMP NORMAL CLEAR MSG RECEIVED WITHOUT ANY ALERT
11838804     ACCEPT MULTIPLE DNS SERVERS FOR ILOM WHERE APPLICABLE IN IPCONF
12357450     FIX TIMESTAMPS IN ALERTS MINED FROM SEL ON V1 SYSTEMS
12708278     CELLSRV SHOULD LOG GD NAME AND OFFSET FOR IO ERRORS
13361797     RS FAILED TO RESTART MS IN RARE SCENARIOS
13495012     ALTER CELL SNMPSUBSCRIBER COMMAND SETS TYPE=ASR ON WRONG ENTRY
13498201     REMOVE SPURIOUS [SERV CELLSRV HANG DETECTED] AFTER CELLSRV CRASH
13521330     NEED TO IMPLEMENT "LIST DATABASE" FOR THE EM IORM UI
13618724     REMOVE ERROR MESSAGE PREFIX FROM ALERT TEXT
13725681     DCLI -K CREATES DIRECTORY /~/.SSH ON SOLARIS 10 NODES.
13737794     VARIABLE-BINDINGS ORDER OF CELLCLI TESTMESSAGE IS NOT VALID.
13807139     IMPROVE RS AND CELLSRV STARTUP IP ERROR REPORTING AND DIAGNOSTICS
13822165     ADD SHOW BANNER AND HIDE STDERR ARGUMENTS TO DCLI
13838283     BETTER ERROR MESSAGE FOR CREATE CELLDISK ON PHYSICAL DISKS IN FAILURE STATUS
13923317     COMMAND PARSING ANOMALY IN CELLCLI RESULTS IN UNEXPECTED SETTINGS
13934957     CELL PATCHING PREREQ-CHECK SHOULD FAIL IF IPCONF -VERIFY IS NOT OK
13934966     CELL PATCHING PREREQ-CHECK SHOULD FAIL IF LIST ALERTHISTORY" SHOW ALERTS
13935080     PATCH SHOULD NOT STATE IT FAILED IF PREREQS ARE NOT OK.
13938302     ALERT IS NOT CREATED IF VALUE OF CL_TEMP METRIC TRESPASSED BUILD-IN THRESHOLD
13973225     MS SHOULD GENERATE ALERTS / TRAPS WHEN SAS LANES IN THE SAS EXPANDER FAIL
13977078     ASR - GRID CONTROL AND ASR TRAP DESTINATION ENTRY REMOVED
14008398     HANDLE INVALID MODEL FROM DMIDECODE MORE GRACEFULLY
14043671     CELLSRV SHUTDOWN WITHOUT FORCE FAILS IF A DISKGROUP IS DISMOUNTED
14045900     SUMMARY TEXT FOR CALIBRATE COMMAND DOES NOT REPORT ERROR 
14065914     FLASH LOGGING FEATURE SHOULD HAVE METRIC FOR BUFFER ALLOCATION FAILURES
14092566     TRACK FLASH CACHE STATE WHEN NOT CACHING
14103957     FLASH CACHE/LOG NOT CREATED ON RESTORED FLASH DISK
14107147     RESYNC TIME SHOULD NOT BE PART OF PATCH TIMEOUT
14148776     ALTER CELLDISK NAME - NEW NAME IS NOT REFLECTED BY LIST FLASHCACHE AND CACHEDBY
14165314     DUPLICATE FS ALERTS SHOULD BE SUPPRESSED
14177001     CUSTOMIZED BATTERY LEARN CYCLE IN EXADATA
14192222     CELL SERVER MODEL SHOULD REFLECT HC OR HP
14199144     MS SENDS ALERTS FROM SNMP TRAP FROM NON-LOCAL SPS
14222004     RENAMING A CELL REGENERATES NEW TEMPERATURE ALERTS AND DONOT CLEAR OLD ALERTS
14239811     MS PD STATUS SHOULD SAY 'FAILED' INSTEAD OF 'CRITICAL'
14244206     MS SCHEDULED BBU RELEARN FAILS TO CHECK BBU STATUS NOR RETRIES 
14263653     NEED PERMANENT FIX FOR 14263651
14305629     DISK INSERTION PROCESSING TOO LONG
14311898     WRONG IORM PLAN OBJECTIVE IS DISPLAYED AFTER DOWNGRADING FROM 11.2.3.2
14312177     SLOW FLASH WITH FLASHLOG ON IT CAUSES CELL SERVICES TO FAIL STARTUP
14313375     ERROR MSG ON /TMP/OC4JPATCH/7439847 : NO SUCH FILE OR DIRECTORY
14356436     ALTER FLASHCACHE FLUSH NEEDS BETTER ERROR MESSAGE FOR WTFC
14366869     CELLSRV DIES WITH SIGSEGV IN FLASHLOGSTORE::SCANACTIVETABLE ALONG LRGSAIOV TEST
14368098     MS: FD WITH CRITICAL PD STATUS STILL SHOWN AS NORMAL
14378866     INCREASE NUMBER OF GRIDDISKS ALLOWED PER BATCH TO AVOID A DOUBLE REBALANCE
14464028     CELLDISK INVALID ERROR WHEN FLASHCACHE CREATED WITH NOT PRESENT STATUS CD 
14480010     FLASH LOG NEEDS IMPROVED CONCURRENCY FOR RE-ENABLING DISKS
14502930     CRITICAL ALERT GENERATED WHEN USB IS REBUILT SUCCESSFULLY 
14505249     IORM OBJECTIVE BALANCED AND LOW_LATENCY DON'T KICK IN FOR SOLO WORKLOAD MODE
14555001     CREATE FLASHCACHE ALL CREATES FLASHCACHE OF SIZE 128M WITH NOT NORMAL CELLDISKS 
14569694     PATCHMGR -CLEANUP SHOULD CLEANING UP _PATCH_HCTAP_/ OR PROVIDE INSTRUCTIONS TO
14588372     ASR - EXADATA FAILING TO SEND SNMP PACKET DUE TO JAVA NULL POINTER ERROR
14610867     FLASH LOG REDO LOG WRITE HISTOGRAM NEEDS TO BE MORE EASILY EXPORTED
14612318     ADD CELLCLI AND SCRIPT SUPPORT TO REPLACE AND RE-ENABLE BBU
14621505     EXADATA GRIDDISK AUTO CREATION FAILED
14646784     JAVA.UTIL.NOSUCHELEMENTEXCEPTION IN MS LOG WHEN EMAIL DELIVERY RETRY FAILS
14674689     FLASH LOG ACTIVE TABLE SIZE SHOULD BE INCREASED
14692944     NEED DETAILED STATS WHY CELL IOS ARE NOT CACHED
14747900     IORM_DATABASE METRICS SHOW DBUA DATABASE METRICS
14758854     FLASHCACHE SIZE CHANGES WHEN CORRUPT CELLDISK IS MADE NORMAL
14769540     DESCRIBE GRIDDISK DOES NOT LIST THE CACHEDBY ATTRIBUTE
14770723     EXPOSE LIFE LEFT ON EACH AURA2 CARD AS A PERCENTAGE 
14803349     DOM CONFINEMENT SHOULD CONSIDER PARTNERSHIP
14808660     CELLSRV FAILED TO START, HIT ORA-600[FLASHLOGPIECELIST::ADDPIECE]
14841844     ASM DOES NOT HAVE ITS OWN DB_* DATABASE METRIC
15871310     CREATE GRIDDISK ALL COMMAND FAILS WHEN A CELLDISK IS NOT NORMAL
15897446     ASR-EXADATA CELL LOCATION OF MIB CAUSING FIFO ERRORS
15963552     POKE FROM CELLSRV IS MISSING WHEN THERE IS NO ASM METADATA
15974057     ALWAYS FLUSH FLASHCACHE FOR WRITEBACK MODE WHEN DOWNGRADING BELOW 11.2.3.2.0
15994904     CELL NEEDS EIGHTH RACK SUPPORT
16001442     FAILED CELLCLI COMMANDS HAVE ZERO EXIT STATUS
16006228     DBSERVER_BACKUP.SH TAR COMMAND SHOULD CORRECTLY HANDLE SPARSE FILES
16028248     LUN AND PHYSICALDISK INFOR NOT CORRECT FOR LIST CELLDISK AFTER RENAME
16064753     ALTER FLASHCACHE ALL WITH A FLASHCACHE ATTRIBUTE REPORTS SUCCESS
16065180     FLASHCACHEMODE SET TO WRITETHROUGH FOR WRB FLASHCACHE
16067726     MS HANG AFTER [OSSMISC:OSSMISC_TIMER_TICKS] WHEN TIME JUMPS BACKWARDS
16074182     ADD "LIST DATABASE" COMMAND TO CELLCLI AND MS
16074653     CPU IMPROVEMENTS FOR FLASHACACHE
16081052     FLASH LOG DISK SIZE SHOULD BE CORRECTLY VALIDATED
16081421     HARD DISK FAILED TO MOUNT AUTOMATICALLY AFTER FLASH DISK FAILURE AND REPLACEMENT
16092303     MS SERVER.XML FILE TRUNCATED AS PART OF A CELL POWER CYCLE
16105593     CELL RPM UPGRADE CHECKING ALL TRACE FILES, CAUSING EXCESSIVE DELAY
16174361     ALTER FLASHCACHE ALL FLUSH DOES NOT FLUSH PEER FAILURE OR POOR PERF CELLDISKS
16193439     ASR/ SNMP TRAP SYSTEM IDENTIFIER FAILURE TO BE SENT ON HALRT FAULTS
16213900     "LIST LUN LUN_NAME" COMMAND FAILS WHEN CELLSRV IS STOPPED
16231174     MS ILLEGALMONITORSTATEEXCEPTION DURING ERASE
16232311     FLASH DISK ALERT SHOULD INCLUDE THE SERIAL NUMBER OF THE WHOLE CARD
16246710     SUNDIAG SHOULD RETRIEVE LSIDIAG_FULL WHEN AURA2.X PROBLEM DETECTED
16278024     FORCE DROP OF WBFC CD GD SHOULD CHECK FOR ASM REDUNDANCY
16278105     REDUCE NUMBER OF CONCURRENT IOS IN A CELL
16371635     CLUSTER WIDE CRASH AFTER ALL CELLSRV CRASHED WITH ORA-600 [KGKPLOALLOC1]
16392070     USE PSID TO UNIQUELY IDENTIFY THE INFINIBAND HCA FOR FIRMWARE CHECK AND UPDATES
16411024     NEED SUPPRESS DISK POWER STATE CHANGE ALERTS
16411732     QUERY/CREATE TABLE ETC. FAILING WITH ORA-27614: SMART I/O FAILED DUE TO AN ERROR
16413066     LIST CELL DOES NOT DISPLAY ALL ATTRIBUTES ON RE-CREATING CELL
16417471     CELLCLI LIST METRICCURRENT FC_IO_ERRS - CELL-02016: METRIC DOES NOT EXIST: FC_IO
16463547     OSS SUPPORT FOR PERMANENT KEEP ACROSS CELLSRV BOUNCE
16472221     WBFC: ORA-600 [FCCGETGDCLS_1] DURING FDOM OFFLINE
16472355     WBFC: CELLSERV HANG DURING FDOM OFFLINE/ONLINE 
16481592     PATCHMGR CLEANUP COLLECTING IRRELEVANT CONTENT
16487249     KERNEL SYSTEM TIME DRIFTING TOO FAST FOR NTP
16495446     UNCORE FREQUENCY IS NOT DISABLED ON X3-2 DB NODES 
16501767     DURING IMAGE UPGRADE TO 11.2.3.2.1 CELL NODE DOES NOT COME UP AFTER LAST REBOOT
16508451     IMPROVE LOG FILE STITCHING BY PATCHMGR
16510225     RE-IMAGE DOESN'T ENABLE ALL CPU CORES
16537444     EXADATA: CELLSRV HANG HAPPENED IN CELLTRANSITION
16538569     CHECKDEVEACHBOOT -FIX MDALL FIX GRUB CAN FAIL WHEN MD4 NOT SYNC'ED
16585329     CONFIGURE LOGROTATE TO ROTATE AND BZIP2 ALL .LOG/.TRC FILES IN /VAR/LOG/CELLOS
16586268     CRASHCORE FILE IS OVERWRITTEN EVERY TIME IN CASE OF OS CRASHDUMP
16590105     FSCK CHECKS NOT DISABLED IN FRESH IMAGE
16591877     ASMDISKGROUPNAME, ASMDISKNAME IS NOT POPULATED CORRECTLY AFTER ASM DG RECREATION
16605828     TEST CASE FOR BUG15882436
16684067     CELLSRV 1M REMOTE RECEIVE PORT BUFFER DEPLETION
16688320     UPON SENDASRTRAP FAILURE, MS FAILED TO SEND REMAINING SNMP SUBSCRIBERS
16688982     WBFC: ASSIGNING FLASH CD IS 10 MINS DELAY AFTER CELLSRV RESTART
16694632     IORM PLAN RESET DOES NOT FREE SUB HEAP EXTENTS
16696321     LNX64-11204-OSS: CELLSRV HITS ORA-600 [DISKIOSCHED::SETPLAN:DB OTHER]
16696985     IMPROVE ALERT LOG MESSAGES WHEN NETWORK ISN'T AVAILABLE
16699385     LNX64-11204-CSS: 160 DB IN ONE CLUSTER, NODE REBOOT AFTER CSSD CRASH
16704019     A "DISK REMOVED" ALERT IS SENT WHEN A DISK ACTUALLY FAILS
16705313     DCLI HANDLES HEAD PIPE INCORRECTLY
16717229     CHANGE DEFAULT VALUE FOR VM/MIN_FREE_KBYTES TO 500MB
16745871     DISABLE THE BUILT-IN CELL AMBIENT THRESHOLD
16768684     FLASHLOG NEEDS PERFORMANCE IMPROVEMENTS BASED ON SLOB RESULTS
16769818     VERIFY-TOPOLOGY NEEDS TO ACCOMODATE NODE_DESC CHANGES & LACK OF SPINE SWITCH
16774368     DISABLE EOIB IN EXADATA IMAGE TO AVOID EXCESSIVE SM LOGGING / LOG WRAP
16775584     MS: "LIST FLASHCACHECONTENT" RETURNS DUPLICATE KEEP OBJECTS
16777412     EXPOSE FLASHCACHE BYPASS REASON METRICS (BUG 14692944 ) VIA MS
16777594     GCW:PATCHMGR CLEANUP DIDN'T HAPPEN WHEN ANY PID MATCH INSTALL.PID
16777751     DCLI DOES NOT CAPTURE REMOTE HOST IDENTIFICATION CHANGED ERROR
16782749     FIX WBFC WRITE METRICS AND ADD REDIRTY METRIC
16796626     TURN OFF OSWATCHER MAKING EXAWATCHER THE ONLY ONE IN USE
16807611     MS FILE DELETION LOGGING AND ROLLING RENAME WRAP ISSUES
16809426     EXADATA ABSOLUTE SERVICE TIME VIOLATION DETECTED ON ONE DISK AFFECTING OTHERS
16815398     REDUCE MTU SIZE ON DB NODES IB INTERFACE TO 7000
16836361     CD_IO_LOAD METRIC VALUES ARE INCORRECT AND TOO HIGH WHEN LOAD IS LOW
16845112     CELLCLI COMMANDS FAIL WITH CELL-2664: FAILED TO CREATE FLASHCACHE ERRORS
16849845     WRB: CELLSRV HANG DURING FDOM FAILURE USING SETPCI
16858835     OUTOFMEMORYERROR RS-7445 [SERV MS NOT RESPONDING] [IT WILL BE RESTARTED]
16864784     TEST NETWORK STATUS
16887059     REMOVE OFED_INFO (RPM OFED-SCRIPTS)
16903390     ORA-600 [PREDICATEDISK::WRITE_5]
16917575     HOTSPARE NOT RECLAIMED WHEN UPGRADING TO 11.2.3.2.1 
16921398     MISSING ARCFOUR CIPHER IN SSHD_CONFIG BREAKS SNAPSHOT
16932116     RESOURCE CONTROL NEEDS MULTIPLE RUNS TO GET CURRENT STATUS 
16949685     TEST DISKGROUP MOUNT WITH ONE INVALID IP IN CELLIP.ORA
16954519     UPDATE COPYRIGHT YEAR 2012 IN CELLCLI BANNER
16964406     DONT LOAD MLX4_EN DRIVER
16973508     WFC: FLASHCACHE SIZE WAS ROUNDED TO MULTIPLE OF 16MB
16977810     SYSTEM DISK IMAGE FAILED DUE TO MD4 NOT DEGRADED
16988043     DISCONTINUE CHECKSWPROFILE.SH
16992011     ALL MDS SHOWS "REMOVED" AFTER REBOOT
16998810     /OPT/ORACLE.EXAWATCHER/ARCHIVE DIRECTORY SHOULD BE ALLOWED TO BE A SYMLINK
17039567     NEED TO UPDATE /ETC/SYSCONFIG/KDUMP KDUMP_COMMANDLINE_APPEND LINE
17084429     RECLAIMDISKS.SH FAILED ABRUPTLY AND -RESTORE OPTION FAILS TO RESTORE.
17088220     ENABLE HWCHECKER IN SOLARIS EXADATA COMPUTE NODE
17157638     PARALLEL DML PRODUCES INCORRECT SQL RESULT
17214800     SUNDIAG SHOULD REFER TO EXAWATCHER INSTEAD OF OSWATCHER
17251471     SPEED UP SINGLE THREAD
17277236     ALTER LUN REENABLE ON FLASH LOG DISK CAN RESULT IN CELLSRV CRASH
17278319     SUNDIAG ENHANCEMENTS FOR ASR, ILOM, EXAWATCHER, CELL CONFINEMENT, NETWORK DATA
17285226     WBFC: CELLSRV HANG IN FLASHCACHECORE.H DUE TO DIRTY LRU QUEUE UPDATE
17295207     REPLACE ALL CURRENT USAGE OF DATE TO DATE FORMAT +'%F %T %Z' FOR LOGS/MESSAGES.
17307247     OSWATCHER NEEDS TO COLLECT MEGARAID FWTERMLOG PERIODICALLY
17313339     SYSTEM DISK REPLACEMENT FAILS ON HP V1 EQUIPMENT
17330822     LNX64-12.1-ASM,CELLSRV CRASH WITH ORA-600[~PREDICATEMAPELEMENT3]
17336036     MISLEADING ERROR ABOUT MISSING XML FILE WHEN UBIOSCONFIG FAILS
17346692     EIGHTH-RACK CONFIGURATION CHANGED TO ENABLED AFTER RESCUE
17349857     NEED TO SET BOOTWITHPINNEDCACHE TO 1 SO THAT EXADATA SYSTEM CAN BOOT WITH PINNED
17362109     ENABLE AUTOMATIC FW UPDATES ON LINUX DB NODES
17371176     DISABLE BUILDS OF NON-UEK OPTIONS FOR DB NODE UPDATE 
17383646     SERIALIZE THE INFINICHECK EXECUTION 
17390553     IORM LIMIT DIRECTIVE CAUSES OVER THROTTLING
17404812     CONCURRENTMODIFICATIONEXCEPTION IN CONFINETRANSITION
17417506     IO ERROR DETAILS FROM CACHE OBJECT
17444979     INFINICHECK FAILS IN EXPANSION STORAGE ONLY MODE
17451210     CELLCLI SERIALIZES CELLMONITOR COMMANDS TOO MUCH
17472203     ADD ALERT DESCRIPTION FOR CONFINED ALERTS; SO EMAILS DO NOT SHOW 'HARDWARE ALERT
17475687     LIST FLASHCACHE CAN DISPLAY VALUES NOT YET POPULATED BY SYNCDISKONCE
17484677     SOME NULL POINTER EXCEPTIONS IN MS
17489799     ROLLBACK FAILED ON V1 BECAUSE OF MISSING USB DEVICES
17504127     EXCESSIVE TRACE ENTRIES WHEN MS RECEIVES A SNMP TRAP
17510981     BLACKLIST EDAC MODULE FROM LOADING
17511671     CALIBRATE FAILED IN CREATING THE RAND_ALL.LUN FILE.
17511684     SET WATCHDOG_THRESH TO 30 AND SET PRINTK TO "4 4 1 7"
17512231     IMPROVE PATCHMGR ERROR MESSAGE WHEN DETECTING ACTIVE ALERTS IN ALERT HISTORY

Tuesday, October 8, 2013

Identify and Diagnostic of Hardware Failure using ILOM on Exadata

To identify Faulty Hardware on Exadata , it use the Sun ILOM.

Below are some steps to Identify and Diagnostics using ILOMs on Exadata Machine using command Lines.

You can Take snapshot and Upload to SR , using ILOM snapshot ASR will help you to Diagnose and provide Further support.

To Take snapshot, Login into ILOM.

Oracle(R) Integrated Lights Out Manager

Version 3.1.2.10 r74387

Copyright (c) 2012, Oracle and/or its affiliates. All rights reserved.

-> help
The help command is used to view information about commands and targets

Usage: help [-format wrap|nowrap] [-o|-output terse|verbose]
[|legal|targets|| ]

Special characters used in the help command are
[]   encloses optional keywords or options
<>   encloses a description of the keyword
     (If <> is not present, an actual keyword is indicated)
|    indicates a choice of keywords or options

help               displays description if this target and its
properties
help     displays description of this property of this target
help targets               displays a list of targets
help legal                 displays the product legal notice

Commands are:
cd
create
delete
dump
exit
help
load
reset
set
show
start
stop
version
--- Choose which type of snapshot you want to Take. 

->set /SP/diag/snapshot dataset=data   [normal|full]
->set /SP/diag/snapshot dump_uri= or   ftp://username:pwd@host_ip_address/~


Identify Hardware Failure.

1) Method.
-> show /SP/faultmgmt

 /SP/faultmgmt
    Targets:
        shell
        0 (/SYS/MB/P0/D7)

2) Method. Which is Very Detailed.

-> show -o table -level all /SP/faultmgmt


Target              | Property               | Value
--------------------+------------------------+---------------------------------
/SP/faultmgmt/0     | fru                    | /SYS/MB/P0/D7
/SP/faultmgmt/0/    | class                  | fault.memory.intel.sb.dimm_ce
 faults/0           |                        |
/SP/faultmgmt/0/    | sunw-msg-id            | SPX86-8004-CE
 faults/0           |                        |
/SP/faultmgmt/0/    | component              | /SYS/MB/P0/D7
 faults/0           |                        |
/SP/faultmgmt/0/    | uuid                   | 34d4bfaa-dummy-ebc8-f95a-dummy-
 faults/0           |                        | d17a
/SP/faultmgmt/0/    | timestamp              | 2013-10-05/23:13:06
 faults/0           |                        |
/SP/faultmgmt/0/    | fru_part_number        | 001-0003
 faults/0           |                        |
/SP/faultmgmt/0/    | fru_dash_level         | 01
 faults/0           |                        |
/SP/faultmgmt/0/    | fru_rev_level          | 50
 faults/0           |                        |
/SP/faultmgmt/0/    | fru_serial_number      | 0000dummy00000dummy
 faults/0           |                        |
/SP/faultmgmt/0/    | fru_manufacturer       | Hynix Semiconductor Inc.
 faults/0           |                        |
/SP/faultmgmt/0/    | fru_name               | 8192MB DDR3 SDRAM DIMM
 faults/0           |                        |
/SP/faultmgmt/0/    | system_manufacturer    | Oracle Corporation
 faults/0           |                        |
/SP/faultmgmt/0/    | system_name            | Exadata X3-2
 faults/0           |                        |
/SP/faultmgmt/0/    | system_part_number     | Exadata X3-2
 faults/0           |                        |
/SP/faultmgmt/0/    | system_serial_number   | AK00122916
 faults/0           |                        |
/SP/faultmgmt/0/    | chassis_manufacturer   | Oracle Corporation
 faults/0           |                        |
/SP/faultmgmt/0/    | chassis_name           | SUN FIRE X4270 M3
 faults/0           |                        |
/SP/faultmgmt/0/    | chassis_part_number    | 700000000
 faults/0           |                        |
/SP/faultmgmt/0/    | chassis_serial_number  | 1323XXXXX03F
 faults/0           |                        |
/SP/faultmgmt/0/    | system_component_manuf | Oracle Corporation
 faults/0           | acturer                |
/SP/faultmgmt/0/    | system_component_name  | SUN FIRE X4270 M3
 faults/0           |                        |
/SP/faultmgmt/0/    | system_component_part_ | 70000000
 faults/0           | number                 |
/SP/faultmgmt/0/    | system_component_seria | 1323XXXXX03F
 faults/0           | l_number               |
/SP/faultmgmt/0/    | serd_count             | 0x7b
 faults/0           |                        |
/SP/faultmgmt/0/    | _list_idx              | 0
 faults/0           |                        |
/SP/faultmgmt/0/    | _list_sz               | 1
 faults/0           |                        | 

Wednesday, April 24, 2013

Incorrect Parallel degree calculated when Auto DOP used in 11.2.0.3 Exadata.


This may be bug in 11.2.0.3 on exadata, AUTO DOP calculate Incorrect Parallel degree, when parallel hints are used with Auto DOP.



SQL> explain plan for
SELECT /*+PARALLEL(E)*/ SITE_ID, ITEM_NUMBER, SERIAL_NUMBER, PORT_NUMBER, EQUIPMENT_ADDRESSABLE, TRIM(STATUS_DATE) 
STATUS_DATE, QUALITY_ASSURANCE_CODE, TRIM(QUALITY_ASSURANCE_DATE) QUALITY_ASSURANCE_DATE, EQUIPMENT_ADDRESS,
EQUIPMENT_OVERRIDE_ACTIVE, INITIALIZE_REQUIRED, PARENTAL_CODE, TEMP_ENABLED, TRIM(TRANSMISSION_DATE) TRANSMISSION_DATE, 
DNS_NAME, IP_ADDRESS, FQDN, LOCAL_STATUS, TRIM( LOCAL_STATUS_DATE) LOCAL_STATUS_DATE, SERVER_ID, SERVER_STATUS, 
TRIM(SERVER_STATUS_DATE) SERVER_STATUS_DATE, EQUIP_DTL_STATUS, PORT_CATEGORY_CODE, HEADEND, ACCOUNT_NUMBER, 
SUB_ACCOUNT_ID, VIDEO_RATING_CODE, PORT_TYPE, SERVICE_CATEGORY_CODE, SERVICE_OCCURRENCE, CREATED_USER_ID, TRIM(DATE_CREATED)
DATE_CREATED, LAST_CHANGE_USER_ID, TRIM(LAST_CHANGE_DATE) LAST_CHANGE_DATE, CABLE_CARD_ID, JOURNAL_DATE 
FROM CLE_DUMMY_DETAIL E 
WHERE JOURNAL_DATE >= (SELECT LAST_RUN_DATE FROM ETL_SITE_RUN_DATE@ODS.WORLD 
WHERE TABLE_NAME = 'xyz' AND SITE_NAME = 'CLE' ) 
AND JOURNAL_DATE <= (SELECT HEART_BEAT_DATE FROM ETL_SITE_RUN_DATE@ODS.WORLD 
WHERE TABLE_NAME = 'xyz' AND SITE_NAME = 'CLE' ) 
;

Explained.

SQL> @xplan

PLAN_TABLE_OUTPUT
----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------
Plan hash value: 350686579

-----------------------------------------------------------------------------------------------------------------------
| Id | Operation | Name | Rows | Bytes | Cost (%CPU)| Time | TQ/Ins |IN-OUT| PQ Distrib |
-----------------------------------------------------------------------------------------------------------------------
| 0 | SELECT STATEMENT | | 1158 | 203K| 18 (0)| 00:00:01 | | | |
| 1 | PX COORDINATOR | | | | | | | | |
| 2 | PX SEND QC (RANDOM)| :TQ10000 | 1158 | 203K| 16 (0)| 00:00:01 | Q1,00 | P->S | QC (RAND) |
| 3 | PX BLOCK ITERATOR |  | 1158 | 203K| 16 (0)| 00:00:01 | Q1,00 | PCWC | |
|* 4 | TABLE ACCESS FULL| CLE_CDO1CPP | 1158 | 203K| 16 (0)| 00:00:01 | Q1,00 | PCWP | |
| 5 | REMOTE | ETL_SITE_RUN_DATE | 1 | 47 | 1 (0)| 00:00:01 | Q1,00 | PCWP | |
| 6 | REMOTE | ETL_SITE_RUN_DATE | 1 | 47 | 1 (0)| 00:00:01 | PODS_~ | R->S | |
-----------------------------------------------------------------------------------------------------------------------

Predicate Information (identified by operation id):
---------------------------------------------------

4 - filter("JOURNAL_DATE">= (SELECT "LAST_RUN_DATE" FROM "ETL_SITE_RUN_DATE" WHERE "SITE_NAME"='CLE' AND
"TABLE_NAME"='xyz') AND "JOURNAL_DATE"<= (SELECT "HEART_BEAT_DATE" FROM "ETL_SITE_RUN_DATE"
WHERE "SITE_NAME"='CLE' AND "TABLE_NAME"='xyz'))

Remote SQL Information (identified by operation id):
----------------------------------------------------

6 - SELECT "SITE_NAME","TABLE_NAME","HEART_BEAT_DATE" FROM "ETL_SITE_RUN_DATE" "ETL_SITE_RUN_DATE" WHERE
"SITE_NAME"='CLE' AND "TABLE_NAME"='xyz' (accessing 'PODS_ETL_CONTROL.WORLD' )


Note
-----
- Degree of Parallelism is 32767 because of hint

31 rows selected.

SQL> show parameter parallel

NAME TYPE VALUE
------------------------------------ ----------- 
parallel_adaptive_multi_user    FALSE
parallel_automatic_tuning      FALSE
parallel_degree_limit       8
parallel_degree_policy      AUTO
parallel_execution_message_size  16384
parallel_force_local       FALSE
parallel_io_cap_enabled      FALSE
parallel_max_servers     196
parallel_min_percent     0
parallel_min_servers     2
parallel_min_time_threshold     60
parallel_server           TRUE
parallel_server_instances    5
parallel_servers_target    90
parallel_threads_per_cpu    2
recovery_parallelism     0

FileName
----------------


Tested with without DBLINK.

-- Testing this on 11.2.0.3

SQL> select /*+ Parallel(E) */ * from foo e;

Execution Plan
----------------------------------------------------------
Plan hash value: 3135133324

--------------------------------------------------------------------------------------------------------------
| Id | Operation | Name | Rows | Bytes | Cost (%CPU)| Time| TQ |IN-OUT| PQ Distrib |
--------------------------------------------------------------------------------------------------------------
| 0 | SELECT STATEMENT | | 18M| 1708M| 9922 (1)| 00:00:01| | | |
| 1 | PX COORDINATOR | | | | || | | |
| 2 | PX SEND QC (RANDOM)| :TQ10000 | 18M| 1708M| 9922 (1)| 00:00:01| Q1,00 | P->S | QC (RAND) |
| 3 | PX BLOCK ITERATOR | | 18M| 1708M| 9922 (1)| 00:00:01| Q1,00 | PCWC | |
| 4 | TABLE ACCESS FULL| FOO | 18M| 1708M| 9922 (1)| 00:00:01| Q1,00 | PCWP | |
--------------------------------------------------------------------------------------------------------------

Note
-----
- Degree of Parallelism is 32767 because of hint

SQL> connect / as sysdba
Connected.
SQL> select count(*) from foo;

COUNT(*)
----------
18472960

parallel_adaptive_multi_user     TRUE
parallel_automatic_tuning      FALSE
parallel_degree_limit       2
parallel_degree_policy      AUTO
parallel_execution_message_size    16384
parallel_force_local       FALSE
parallel_instance_group 
parallel_io_cap_enabled      FALSE
parallel_max_servers       135
parallel_min_percent       0
parallel_min_servers       0
parallel_min_time_threshold     1
parallel_server        FALSE
parallel_server_instances      1
parallel_servers_target      64
parallel_threads_per_cpu      2
recovery_parallelism       0



comparison on 11.2.0.2 and 11.2.0.3

----test case 
drop table foo;
alter session set parallel_degree_policy = auto;
alter session set parallel_degree_limit = 2;
alter session set parallel_min_time_threshold = 1;
create table foo as select * from all_objects where rownum < 16;
alter table foo parallel 6;
set autotrace traceonly explain
select /*+ Parallel(e) */ * from foo e;

--------
FOO has a default DOP of 6

11.2.0.3 -- select /*+ Parallel(a) */ * from foo a; --- Degree of Parallelism is 32767 because of hint
11.2.0.2 -- select /*+ Parallel(a) */ * from foo a; -- automatic DOP: Computed Degree of Parallelism is 1
11.2.0.3 select /*+ Parallel */ * from foo; -- - automatic DOP: Computed Degree of Parallelism is 2
11.2.0.2 select /*+ Parallel */ * from foo; -- - automatic DOP: Computed Degree of Parallelism is 2
11.2.0.3 select /*+ Parallel(20) */ * from foo; -- - Degree of Parallelism is 20 because of hint
11.2.0.2 select /*+ Parallel(20) */ * from foo; -- - Degree of Parallelism is 20 because of hint


Friday, December 21, 2012

Exadata Backup on ZFS.

Below is Test of Exadata x2 half rack Database backup on ZFS.


Size : 2 TB
Used Size : 1.6 TB
Time Taken: 1hr 34 Mins.
Storage: 4 share of 300TB ZFS.
Channels Allocated: 5 channels per node per share. Totally 20 channels across 4 nodes and 4 shares.


#!/bin/ksh
export ORACLE_HOME=/u01/app/oracle/product/11.2.0.3/db_1
export ORACLE_SID=DUMMYDB1
export PATH=$ORACLE_HOME/bin:$PATH
export TNS_ADMIN=/u01/app/oracle/product/11.2.0.3/db_1/network/admin
export ALTER SESSION NLS_DATE_FORMAT='DD-MON-YYYY HH24:MI'
export LOGFILE=/home/oracle/DUMMYDB_rman_zfs_level0_ver3.log

#setup oracle environment based on SID
ORAENV_ASK=NO
. oraenv

rman target / catalog rman/xxxx@catalogdb <> ${LOGFILE}
run {
sql 'alter system set "_backup_disk_bufcnt"=64 scope=memory';
sql 'alter system set "_backup_disk_bufsz"=1048576 scope=memory';
allocate channel c01 DEVICE TYPE DISK FORMAT '/export/share8/DUMMYDB/DUMMYDBtest301_%U' CONNECT 'sys/XXXX@DUMMYDB1.WORLD';
allocate channel c02 DEVICE TYPE DISK FORMAT '/export/share8/DUMMYDB/DUMMYDBtest302_%U' CONNECT 'sys/XXXX@DUMMYDB1.WORLD';
allocate channel c03 DEVICE TYPE DISK FORMAT '/export/share8/DUMMYDB/DUMMYDBtest303_%U' CONNECT 'sys/XXXX@DUMMYDB1.WORLD';
allocate channel c04 DEVICE TYPE DISK FORMAT '/export/share8/DUMMYDB/DUMMYDBtest304_%U' CONNECT 'sys/XXXX@DUMMYDB1.WORLD';
allocate channel c05 DEVICE TYPE DISK FORMAT '/export/share8/DUMMYDB/DUMMYDBtest305_%U' CONNECT 'sys/XXXX@DUMMYDB1.WORLD';
allocate channel c06 DEVICE TYPE DISK FORMAT '/export/share7/DUMMYDB/DUMMYDBtest306_%U' CONNECT 'sys/XXXX@DUMMYDB3.WORLD';
allocate channel c07 DEVICE TYPE DISK FORMAT '/export/share7/DUMMYDB/DUMMYDBtest307_%U' CONNECT 'sys/XXXX@DUMMYDB3.WORLD';
allocate channel c08 DEVICE TYPE DISK FORMAT '/export/share7/DUMMYDB/DUMMYDBtest308_%U' CONNECT 'sys/XXXX@DUMMYDB3.WORLD';
allocate channel c09 DEVICE TYPE DISK FORMAT '/export/share7/DUMMYDB/DUMMYDBtest309_%U' CONNECT 'sys/XXXX@DUMMYDB3.WORLD';
allocate channel c10 DEVICE TYPE DISK FORMAT '/export/share7/DUMMYDB/DUMMYDBtest310_%U' CONNECT 'sys/XXXX@DUMMYDB3.WORLD';
allocate channel c11 DEVICE TYPE DISK FORMAT '/export/share6/DUMMYDB/DUMMYDBtest311_%U' CONNECT 'sys/XXXX@DUMMYDB2.WORLD';
allocate channel c12 DEVICE TYPE DISK FORMAT '/export/share6/DUMMYDB/DUMMYDBtest312_%U' CONNECT 'sys/XXXX@DUMMYDB2.WORLD';
allocate channel c13 DEVICE TYPE DISK FORMAT '/export/share6/DUMMYDB/DUMMYDBtest313_%U' CONNECT 'sys/XXXX@DUMMYDB2.WORLD';
allocate channel c14 DEVICE TYPE DISK FORMAT '/export/share6/DUMMYDB/DUMMYDBtest314_%U' CONNECT 'sys/XXXX@DUMMYDB2.WORLD';
allocate channel c15 DEVICE TYPE DISK FORMAT '/export/share6/DUMMYDB/DUMMYDBtest315_%U' CONNECT 'sys/XXXX@DUMMYDB2.WORLD';
allocate channel c16 DEVICE TYPE DISK FORMAT '/export/share5/DUMMYDB/DUMMYDBtest316_%U' CONNECT 'sys/XXXX@DUMMYDB4.WORLD';
allocate channel c17 DEVICE TYPE DISK FORMAT '/export/share5/DUMMYDB/DUMMYDBtest317_%U' CONNECT 'sys/XXXX@DUMMYDB4.WORLD';
allocate channel c18 DEVICE TYPE DISK FORMAT '/export/share5/DUMMYDB/DUMMYDBtest318_%U' CONNECT 'sys/XXXX@DUMMYDB4.WORLD';
allocate channel c19 DEVICE TYPE DISK FORMAT '/export/share5/DUMMYDB/DUMMYDBtest319_%U' CONNECT 'sys/XXXX@DUMMYDB4.WORLD';
allocate channel c20 DEVICE TYPE DISK FORMAT '/export/share5/DUMMYDB/DUMMYDBtest320_%U' CONNECT 'sys/XXXX@DUMMYDB4.WORLD';
sql "alter system archive log current";
backup incremental level 0 database plus archivelog;
sql "alter database backup controlfile to trace";
release channel c01;
release channel c02;
release channel c03;
release channel c04;
release channel c05;
release channel c06;
release channel c07;
release channel c08;
release channel c09;
release channel c10;
release channel c11;
release channel c12;
release channel c13;
release channel c14;
release channel c15;
release channel c16;
release channel c17;
release channel c18;
release channel c19;
release channel c20;
}
exit
EOF

References:

http://www.oracle.com/technetwork/database/features/availability/maa-wp-dbm-zfs-backup-1593252.pdf

http://www.oracle.com/technetwork/database/features/availability/maa-tech-wp-sundbm-backup-11202-183503.pdf

Monday, October 22, 2012

Exadata Upgrade / Upgrade GI to 11.2.0.3

Few days back we did upgrade from 11.2.0.2 to 11.2.0.3 on Exadata x2 half rack.

This was direct upgrade from 11.2.0.2 to 11.2.0.3 with Bundle Patch 7  /  patch 13992240

Some High level steps. 
  • Oracle Grid Infrastructure 11.2.0.3 Installation (Dont run rootupgrade.sh at the end)
  • Install BP7 on New Grid home.
  • Run rootupgrade.sh (Outage Time)
  • 11.2.0.3 database software Installation. 
  • Install BP7 on New Oracle home
  • Upgrade database to 11.2.0.3 (Outage Time)

Attached is Detail Upgrade Document on Google docs.



Let me know if you any difficulties viewing above Document on jonyjt@gmail.com



Tuesday, August 14, 2012

resmgr:pq queued | enq: JX - SQL statement queue | PX Queuing: statement queue

We have 4 Node RAC on 11.2.0.2 on exadata machine.

Few days back we had problem of Trucate statment was running slow. it took around 1 hour.
when truncate is running slow.
either object is locked or its bug. But here it was something else.

TRUNCATE TABLE DUMMY_USAGE.DUMMY_CUST_BILL_CYC_USG_F_BLD1 DROP STORAGE;

After pulling AWR , we found that "resmgr:pq queued" is in Top 5 wait events.
It looks that database has been suffered from "lack of enough parallel servers".
even TRUNCATE statement also tried to ran into parallel.


in 11.2.0.2

resmgr:pq queued is The time the session waited for sufficient parallel query processes to become available to run this session with the requested degree of parallelism

in 11.2.0.1

Wait event "PX Queuing: statement queue" is the event When statement waits on about to run.

Wait event "enq: JX - SQL statement queue" is event when statement have few more statments lined up ahead of it.


Wait event on DB. So database has heavily waited on "resmgr:pq queued"

MIN(SAMPLE_TIME) MAX(SAMPLE_TIME)                 COUNT EVENT
8/10/2012 3:39:15.324 AM 8/10/2012 7:00:26.064 AM 11812 resmgr:pq queued
8/10/2012 2:13:36.850 AM 8/10/2012 7:00:16.054 AM 3045 CPU
8/10/2012 3:05:19.667 AM 8/10/2012 6:59:03.635 AM 2638 direct path read temp
8/10/2012 2:00:45.579 AM 8/10/2012 6:59:15.940 AM 1987 cell smart table scan

Which SQL has waited Most from "lack of Enogh parallel servers"

SELECT MIN(SAMPLE_TIME),MAX(SAMPLE_TIME),COUNT(*) AS COUNT , EVENT,SQL_ID FROM DBA_HIST_ACTIVE_sESS_HISTORY WHERE SNAP_ID BETWEEN 1594 AND 1598 AND EVENT='resmgr:pq queued' group by EVENT,SQL_ID ORDER BY COUNT DESC;

MIN(SAMPLE_TIME)         MAX(SAMPLE_TIME)        COUNT    EVENT           SQL_ID
---------------------------------------------------------------------------------------------
8/10/2012 4:24:47.633 AM 8/10/2012 6:25:40.210 AM 723 resmgr:pq queued brjbmruuj4qca

8/10/2012 4:06:58.060 AM 8/10/2012 5:49:08.703 AM 540 resmgr:pq queued 1ztm7k9fbawks
8/10/2012 4:07:05.815 AM 8/10/2012 5:38:45.312 AM 477 resmgr:pq queued 8fdq282ydmj9y
8/10/2012 4:06:58.060 AM 8/10/2012 5:37:57.564 AM 473 resmgr:pq queued 5pgx9fqn6b18k
8/10/2012 4:06:58.060 AM 8/10/2012 5:37:07.457 AM 470 resmgr:pq queued bfyb7a00r9m97
8/10/2012 4:06:58.060 AM 8/10/2012 5:34:07.155 AM 452 resmgr:pq queued dck4f08ywn2gk
8/10/2012 4:07:05.815 AM 8/10/2012 5:33:34.783 AM 448 resmgr:pq queued 2tjdy0su2yyrg
8/10/2012 4:06:58.060 AM 8/10/2012 5:32:26.963 AM 442 resmgr:pq queued a770bk5gnyp49
8/10/2012 4:06:58.060 AM 8/10/2012 5:31:56.913 AM 439 resmgr:pq queued 8qgm5az2pstpw
8/10/2012 4:07:05.815 AM 8/10/2012 5:31:24.561 AM 435 resmgr:pq queued 3gytqk6h6hy46
8/10/2012 4:07:05.815 AM 8/10/2012 5:31:14.513 AM 434 resmgr:pq queued 73bxznwgh04yf

8/10/2012 4:49:52.568 AM 8/10/2012 5:49:08.703 AM 356 resmgr:pq queued 3qkqq4w2z03jh
8/10/2012 4:06:58.060 AM 8/10/2012 5:03:13.976 AM 338 resmgr:pq queued g0b4p5zsn5xwp
8/10/2012 4:06:58.060 AM 8/10/2012 5:02:23.885 AM 333 resmgr:pq queued 0k5796sr9sgrq
8/10/2012 4:07:05.815 AM 8/10/2012 5:01:01.402 AM 324 resmgr:pq queued a8m336ddng5z7
8/10/2012 4:07:05.815 AM 8/10/2012 4:59:41.282 AM 316 resmgr:pq queued 8wz2qpu6t3b4r
8/10/2012 4:06:58.060 AM 8/10/2012 4:58:53.520 AM 312 resmgr:pq queued c1a03akcb57nm
8/10/2012 4:39:51.493 AM 8/10/2012 5:31:06.810 AM 308 resmgr:pq queued bsuwb8ksjgg7f
8/10/2012 5:38:15.272 AM 8/10/2012 6:32:10.874 AM 287 resmgr:pq queued 8tkfm8yg6fy3z

SQL> SELECT * FROM DBA_HIST_SQLTEXT WHERE sql_id in ('brjbmruuj4qca','1ztm7k9fbawks');

      DBID SQL_ID        SQL_TEXT                                                                         COMMAND_TYPE
---------- ------------- -------------------------------------------------------------------------------- ------------
1664458898 brjbmruuj4qca TRUNCATE TABLE DUMMY_USAGE.DUMMY_TOP_TALKERS_DLY_SUM_SWP DROP STORAGE                        85
2309640764 8fdq282ydmj9y SELECT Last_day(cduf.time_key)         AS time_key                                          3


when Oracle put statments to Queue for parallel servers.

When the parameter PARALLEL_DEGREE_POLICY is set to AUTO, Oracle Database queues SQL statements that require parallel execution if the necessary parallel server processes are not available. After the necessary resources become available, the SQL statement is dequeued and allowed to execute. The default dequeue order is a simple first in, first out queue based on the time a statement was issued.

The following is a summary of parallel statement processing.

1) A SQL statements is issued.

2) The statement is parsed and the DOP is automatically determined.

3) Available parallel resources are checked.

A) If there are enough parallel resources and there are no statements ahead in the queue waiting for the resources, the SQL statement is executed.

B) If there are not enough parallel servers, the SQL statement is queued based on specified conditions and dequeued from the front of the queue when specified conditions are met.

Parallel statements are queued if running the statements would increase the number of active parallel servers above the value of the PARALLEL_SERVERS_TARGET initialization parameter. For example, if PARALLEL_SERVERS_TARGET is set to 64, the number of current active servers is 60, and a new parallel statement needs 16 parallel servers, it would be queued because 16 added to 60 is greater than 64, the value of PARALLEL_SERVERS_TARGET.

The default value is described in "PARALLEL_SERVERS_TARGET". This value is not the maximum number of parallel server processes allowed on the system, but the number available to run parallel statements before parallel statement queuing is used. It is set lower than the maximum number of parallel server processes allowed on the system (PARALLEL_MAX_SERVERS) to ensure each parallel statement gets all of the parallel server resources required and to prevent overloading the system with parallel server processes. Note all serial (nonparallel) statements execute immediately even if parallel statement queuing has been activated.

If a statement has been queued, it is identified by the resmgr:pq queued wait event.

Parameter to define parallelism were defined something like below,
SQL> show parameter parallel

NAME                                 TYPE        VALUE
------------------------------------ ----------- ------------------------------
fast_start_parallel_rollback         string      LOW
parallel_adaptive_multi_user         boolean     FALSE
parallel_automatic_tuning            boolean     FALSE
parallel_degree_limit                string      16
parallel_degree_policy               string      AUTO   --AUTO DOP is used. 
parallel_execution_message_size      integer     16384
parallel_force_local                 boolean     FALSE
parallel_instance_group              string
parallel_io_cap_enabled              boolean     FALSE
parallel_max_servers                 integer     32
parallel_min_time_threshold          string      AUTO
parallel_server                      boolean     TRUE
parallel_server_instances            integer     4   
parallel_servers_target              integer     16
parallel_threads_per_cpu             integer     1

SQL> show parameter cpu

NAME                                 TYPE        VALUE
------------------------------------ ----------- ------------------------------
cpu_count                            integer     24  
parallel_threads_per_cpu             integer     1


It have only 32 Parallel servers and queuing will start even if 16 servers are Used.

so Both parallel_max_servers & parallel_servers_target have set incorrect.

Auto DOP is calculated as (parallel_threads_per_cpu × parallel_server_instances × cpu_count )= 96

The default value for parallel_server_target is set to 4 times the default DOP.

((4 × CPU_count) × parallel_threads_per_cpu) × active_instances = PARALLEL_SERVER_TARGET

((4 × 24) × 1) × 4) = 384 should be PARALLEL_SERVER_TARGET


Managing Parallel Statement Queuing with Hints

NO_STATEMENT_QUEUING

When PARALLEL_DEGREE_POLICY is set to AUTO, this hint enables a statement to bypass the parallel statement queue. For example:

SELECT /*+ NO_STATEMENT_QUEUING */ emp.last_name, dpt.department_name
FROM employees emp, departments dpt
WHERE emp.department_id = dpt.department_id;

STATEMENT_QUEUING

When PARALLEL_DEGREE_POLICY is not set to AUTO, this hint enables a statement to be delayed and to only run when parallel processes are available to run at the requested DOP. For example:

SELECT /*+ STATEMENT_QUEUING */ emp.last_name, dpt.department_name
FROM employees emp, departments dpt
WHERE emp.department_id = dpt.department_id;

There is also a hidden parameter which control parallel queuing.
_parallel_statement_queuing=TRUE





Friday, July 27, 2012

Configure Parallel DEGREE / INSTANCES of Table / PARALLEL execution running slow.

This is demo about Query which performing Parallel but takes lot of time.
Configuring Table parallel degree / instances degree incorrect can result adverse effect.

Platform is Exadata-X2.

Elapsed time has dropped from 2 minutes to 20 seconds.

SELECT count(1)
FROM DUMMY_TPLGY
WHERE (DUMMY_TPLGY.DVIC_TYPE = 'optical_node'
OR DUMMY_TPLGY.DVIC_TYPE = 'amplifier')
AND DUMMY_TPLGY.SNPSHT_DTM in
(SELECT SNPSHT_DTM
FROM DUMMY_TPLGY_CTRL
WHERE AVL_IND = 'Y'
AND SNPSHT_DTM >
(SELECT MAX (DAY_KEY)
FROM RCM_CTRL
WHERE TBL_NAME = 'DUMMY_DIM' AND AVL_IND = 'Y'));

Lets see What explain plan says .

set timing on 
set time on 
set long 999999
set lines 180
set pages 555

12:24:11 SQL> explain plan for
12:24:11   2  SELECT count(1)
12:24:11   3  FROM DUMMY_TPLGY
12:24:11   4  WHERE (DUMMY_TPLGY.DVIC_TYPE = 'optical_node'
12:24:11   5  OR DUMMY_TPLGY.DVIC_TYPE = 'amplifier')
12:24:11   6  AND DUMMY_TPLGY.SNPSHT_DTM in
12:24:11   7  (SELECT SNPSHT_DTM
12:24:11   8  FROM DUMMY_TPLGY_CTRL
12:24:11   9  WHERE AVL_IND = 'Y'
12:24:11  10  AND SNPSHT_DTM >
12:24:11  11  (SELECT MAX (DAY_KEY)
12:24:11  12  FROM RCM_CTRL
12:24:11  13  WHERE TBL_NAME = 'DUMMY_DIM' AND AVL_IND = 'Y'));


Explained.

Elapsed: 00:00:02.05

As you can see cost are too much high for full table scan.

12:24:15 SQL> select * from table(dbms_xplan.display);

PLAN_TABLE_OUTPUT
------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------
Plan hash value: 231818000

--------------------------------------------------------------------------------------------------------------------------------------------------
| Id  | Operation                     | Name                | Rows  | Bytes | Cost (%CPU)| Time     | Pstart| Pstop |    TQ  |IN-OUT| PQ Distrib |
--------------------------------------------------------------------------------------------------------------------------------------------------
|   0 | SELECT STATEMENT              |                     |     1 |    26 | 50332   (1)| 00:10:04 |       |       |        |         |          |
|   1 |  SORT AGGREGATE               |                     |     1 |    26 |            |          |       |       |        |         |          |
|   2 |   NESTED LOOPS                |                     |  3434K|    85M| 49646   (1)| 00:09:56 |       |       |        |         |          |
|*  3 |    TABLE ACCESS BY INDEX ROWID| DUMMY_TPLGY_CTRL     |     3 |    30 |     1   (0)| 00:00:01 |       |       |        |         |          |
|*  4 |     INDEX RANGE SCAN          | XPK_DUMMY_TPLGY_CTRL |     1 |       |     1   (0)| 00:00:01 |       |       |        |         |          |
|   5 |      SORT AGGREGATE           |                     |     1 |    26 |            |          |       |       |        |         |          |
|   6 |       PX COORDINATOR          |                     |       |       |            |          |       |       |        |         |          |
|   7 |        PX SEND QC (RANDOM)    | :TQ10000            |     1 |    26 |            |          |       |       |  Q1,00 |    P->S | QC (RAND)|
|   8 |         SORT AGGREGATE        |                     |     1 |    26 |            |          |       |       |  Q1,00 |    PCWP |          |
|   9 |          PX BLOCK ITERATOR    |                     |   382 |  9932 |    24   (0)| 00:00:01 |       |       |  Q1,00 |    PCWC |          |
|* 10 |           TABLE ACCESS FULL   | RCM_CTRL            |   382 |  9932 |    24   (0)| 00:00:01 |       |       |  Q1,00 |    PCWP |          |
|  11 |    PARTITION LIST ITERATOR    |                     |  1174K|    17M| 16548   (1)| 00:03:19 |   KEY |   KEY |        |         |          |
|* 12 |     TABLE ACCESS FULL         | DUMMY_TPLGY          |  1174K|    17M| 16548   (1)| 00:03:19 |   KEY |   KEY |        |         |          |
--------------------------------------------------------------------------------------------------------------------------------------------------

Predicate Information (identified by operation id):
---------------------------------------------------

   3 - filter("AVL_IND"='Y')
   4 - access("SNPSHT_DTM"> (SELECT MAX(SYS_OP_CSR(SYS_OP_MSR(MAX("DAY_KEY")),0)) FROM "RCM"."RCM_CTRL" "RCM_CTRL" WHERE
              "TBL_NAME"='DUMMY_DIM' AND "AVL_IND"='Y'))
  10 - filter("TBL_NAME"='DUMMY_DIM' AND "AVL_IND"='Y')
  12 - filter(("DUMMY_TPLGY"."DVIC_TYPE"='amplifier' OR "DUMMY_TPLGY"."DVIC_TYPE"='optical_node') AND
              "DUMMY_TPLGY"."SNPSHT_DTM"="SNPSHT_DTM")

Note
-----
   - dynamic sampling used for this statement (level=6)

33 rows selected.

Lets see actual execution Time .

Elapsed: 00:00:01.01
12:24:28 SQL>
12:24:28 SQL>
12:28:14 SQL> SELECT count(1)
12:28:15   2  FROM DUMMY_TPLGY
12:28:15   3  WHERE (DUMMY_TPLGY.DVIC_TYPE = 'optical_node'
12:28:15   4  OR DUMMY_TPLGY.DVIC_TYPE = 'amplifier')
12:28:15   5  AND DUMMY_TPLGY.SNPSHT_DTM in
12:28:15   6  (SELECT SNPSHT_DTM
12:28:15   7  FROM DUMMY_TPLGY_CTRL
12:28:15   8  WHERE AVL_IND = 'Y'
12:28:15   9  AND SNPSHT_DTM >
12:28:15  10  (SELECT MAX (DAY_KEY)
12:28:15  11  FROM RCM_CTRL
12:28:15  12  WHERE TBL_NAME = 'DUMMY_DIM' AND AVL_IND = 'Y'));

  COUNT(1)
----------
   3574152

Elapsed: 00:02:07.34
12:30:23 SQL>

Almost 2 minutes.

What are these tables and why it takes so long even if its running PARALLEL.
Culprit is Table DUMMY_TPLGY, which is going for FULL TABLE SCAN and is very big.

SQL> SELECT OWNER,TABLE_NAME,DEGREE,INSTANCES,PARTITIONED,NUM_ROWS FROM DBA_TABLES WHERE TABLE_NAME IN ('DUMMY_TPLGY','RCM_CTRL','DUMMY_TPLGY_CTRL');

OWNER      TABLE_NAME           DEGREE     INSTANCES  PARTITIONED       NUM_ROWS
---------- -------------------- ---------- ---------- --------------- ----------
RCM        RCM_CTRL                DEFAULT    DEFAULT NO                    4597
TNC        DUMMY_TPLGY                    1          1 YES              277101036
TNC        DUMMY_TPLGY_CTRL               1          1 NO                      59

----------Now Tables are in PARALLEL but degree is one Only.
So lets change it default and let oracle decide how to execute it.

SQL> ALTER TABLE TNC.DUMMY_TPLGY PARALLEL (DEGREE DEFAULT INSTANCES DEFAULT);

Table altered.

SQL> ALTER TABLE TNC.DUMMY_TPLGY_CTRL PARALLEL (DEGREE DEFAULT INSTANCES DEFAULT);

Table altered.

SQL> SELECT OWNER,TABLE_NAME,DEGREE,INSTANCES,PARTITIONED,NUM_ROWS FROM DBA_TABLES WHERE TABLE_NAME IN ('DUMMY_TPLGY','RCM_CTRL','DUMMY_TPLGY_CTRL');

OWNER      TABLE_NAME           DEGREE     INSTANCES  PARTITIONED       NUM_ROWS
---------- -------------------- ---------- ---------- --------------- ----------
RCM        RCM_CTRL                DEFAULT    DEFAULT NO                    4597
TNC        DUMMY_TPLGY              DEFAULT    DEFAULT YES              277101036
TNC        DUMMY_TPLGY_CTRL         DEFAULT    DEFAULT NO                      59

---New execution plan has low cost as compare to previous plan even Execution time also dropped drastically .

12:34:38 SQL> select * from table(dbms_xplan.display);

PLAN_TABLE_OUTPUT
------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------
Plan hash value: 3873636908

--------------------------------------------------------------------------------------------------------------------------------------------------------
| Id  | Operation                           | Name                | Rows  | Bytes | Cost (%CPU)| Time     | Pstart| Pstop |    TQ  |IN-OUT| PQ Distrib|
-------------------------------------------------------------------------------------------------------------------------------------------------------
|   0 | SELECT STATEMENT                    |                     |     1 |    26 |  1746   (1)| 00:00:21 |       |       |        |      |           |
|   1 |  SORT AGGREGATE                     |                     |     1 |    26 |            |          |       |       |        |      |           |
|   2 |   PX COORDINATOR                    |                     |       |       |            |          |       |       |        |      |           |
|   3 |    PX SEND QC (RANDOM)              | :TQ20001            |     1 |    26 |            |          |       |       |  Q2,01 | P->S | QC (RAND) |
|   4 |     SORT AGGREGATE                  |                     |     1 |    26 |            |          |       |       |  Q2,01 | PCWP |           |
|   5 |      NESTED LOOPS                   |                     |  3434K|    85M|  1723   (1)| 00:00:21 |       |       |  Q2,01 | PCWP |           |
|   6 |       BUFFER SORT                   |                     |       |       |            |          |       |       |  Q2,01 | PCWC |           |
|   7 |        PX RECEIVE                   |                     |       |       |            |          |       |       |  Q2,01 | PCWP |           |
|   8 |         PX SEND BROADCAST           | :TQ20000            |       |       |            |          |       |       |        | S->P | BROADCAST |
|*  9 |          TABLE ACCESS BY INDEX ROWID| DUMMY_TPLGY_CTRL     |     3 |    30 |     1   (0)| 00:00:01 |       |       |        |      |           |
|* 10 |           INDEX RANGE SCAN          | XPK_DUMMY_TPLGY_CTRL |     1 |       |     1   (0)| 00:00:01 |       |       |        |      |           |
|  11 |            SORT AGGREGATE           |                     |     1 |    26 |            |          |       |       |        |      |           |
|  12 |             PX COORDINATOR          |                     |       |       |            |          |       |       |        |      |           |
|  13 |              PX SEND QC (RANDOM)    | :TQ10000            |     1 |    26 |            |          |       |       |Q1,00   | P->S | QC (RAND) |
|  14 |               SORT AGGREGATE        |                     |     1 |    26 |            |          |       |       |Q1,00   | PCWP |           |
|  15 |                PX BLOCK ITERATOR    |                     |   382 |  9932 |    24   (0)| 00:00:01 |       |       |Q1,00   | PCWC |           |
|* 16 |                 TABLE ACCESS FULL   | RCM_CTRL            |   382 |  9932 |    24   (0)| 00:00:01 |       |       |Q1,00   | PCWP |           |
|  17 |       PX BLOCK ITERATOR             |                     |  1174K|    17M|   574   (1)| 00:00:07 |   KEY |   KEY |Q2,01   | PCWC |           |
|* 18 |        TABLE ACCESS FULL            | DUMMY_TPLGY          |  1174K|    17M|   574   (1)| 00:00:07 |   KEY |   KEY |Q2,01   | PCWP |           |
-------------------------------------------------------------------------------------------------------------------------------------------------------

Predicate Information (identified by operation id):
---------------------------------------------------

   9 - filter("AVL_IND"='Y')
  10 - access("SNPSHT_DTM"> (SELECT MAX(SYS_OP_CSR(SYS_OP_MSR(MAX("DAY_KEY")),0)) FROM "RCM"."RCM_CTRL" "RCM_CTRL" WHERE "TBL_NAME"='DUMMY_DIM'
              AND "AVL_IND"='Y'))
  16 - filter("TBL_NAME"='DUMMY_DIM' AND "AVL_IND"='Y')
  18 - filter(("DUMMY_TPLGY"."DVIC_TYPE"='amplifier' OR "DUMMY_TPLGY"."DVIC_TYPE"='optical_node') AND "DUMMY_TPLGY"."SNPSHT_DTM"="SNPSHT_DTM")

---Now Lets run and Measure execution time.

SQL> ALTER SESSION ENABLE PARALLEL QUERY;

Session altered.

SQL> SET TIME ON
12:51:00 SQL> SET TIMING ON;
12:51:03 SQL>
12:51:15 SQL> SELECT count(1)
12:51:16   2  FROM DUMMY_TPLGY
12:51:16   3  WHERE (DUMMY_TPLGY.DVIC_TYPE = 'optical_node'
12:51:16   4  OR DUMMY_TPLGY.DVIC_TYPE = 'amplifier')
12:51:16   5  AND DUMMY_TPLGY.SNPSHT_DTM in
12:51:16   6  (SELECT SNPSHT_DTM
12:51:16   7  FROM DUMMY_TPLGY_CTRL
12:51:16   8  WHERE AVL_IND = 'Y'
12:51:16   9  AND SNPSHT_DTM >
12:51:16  10  (SELECT MAX (DAY_KEY)
12:51:16  11  FROM RCM_CTRL
12:51:16  12  WHERE TBL_NAME = 'DUMMY_DIM' AND AVL_IND = 'Y'));

  COUNT(1)
----------
   3574152

Elapsed: 00:00:20.89
12:51:37 SQL>


Here we go only 20 seconds from 2 minutes.